December 2024 Summaries
4 posts from Together AI
Filter
Month:
Year:
Post Summaries
Back to Blog
Serverless LoRA inference with pay-per-token pricing allows users to upload their own LoRA adapters and run inference on them alongside a compatible serverless model, including popular models like Llama 3.1 and Qwen 2.5. The platform enables dynamic adapter switching at scale, running hundreds of models for the same price as a single base model. This allows for cost-efficient model customization, faster iteration and experimentation, optimized performance at scale, and easy fine-tuning of custom LoRA adapters with the Together Fine-tuning API.
Dec 18, 2024
1,224 words in the original blog post.
Together AI has acquired CodeSandbox, a development environment infrastructure used by over 4.5 million developers monthly, to launch its first-of-its-kind code interpreter for generative AI. This technology enables LLMs to execute generated code, allowing developers to harness the full potential of LLMs and build more complex applications. The Together Code Interpreter has already generated excitement among developers at AI-native companies, including Blackbox AI, which relies on CodeSandbox's SDK for its coding agents. CodeSandbox is also launching a new SDK beta that enables developers to create and run VM sandboxes quickly and securely. This acquisition marks a significant milestone in making generative AI more accessible, versatile, and impactful for developers.
Dec 12, 2024
932 words in the original blog post.
Together AI has partnered with Meta to support the latest advancement in the Llama model series, Llama-3.3-70B-Instruct. The new Llama 3.3-70B model offers enhanced reasoning, mathematics, and instruction-following capabilities comparable to the much larger Llama 3.1 405B model at a fraction of the cost. Over 250,000 AI developers and enterprises are leveraging Together AI Inference and Fine-Tuning Platform for innovation with AI. To support Llama 3.3-70B, Together AI is introducing a Turbo serverless endpoint that offers unparalleled performance, accuracy, and affordability. Dedicated Endpoints will also be offered soon, providing consistent quality unaffected by other users' load. The Llama 3.3 model is part of the open-source ecosystem, offering transparency, flexibility, and community-driven innovation.
Dec 06, 2024
500 words in the original blog post.
Amazon Web Services (AWS) Marketplace now offers Together AI, a platform designed to accelerate enterprise AI development. With access to over 200 open-source models, including Llama-3.1 and Qwen 2.5, users can experience up to 2-3x faster inference times and optimized performance. AWS customers can utilize their existing cloud credits and committed spend while cutting GPU costs by up to a third compared to closed-source solutions. Together AI's collaboration with AWS aims to accelerate enterprise generative AI development through enhanced inference speed, reduced GPU consumption, and improved model orchestration capabilities.
Dec 02, 2024
415 words in the original blog post.