February 2025 Summaries
4 posts from Together AI
Filter
Month:
Year:
Post Summaries
Back to Blog
Minions, a method that collaborates between small on-device models and frontier cloud models, reduces cloud costs while maintaining performance. Small LMs are improving rapidly and can tackle real tasks, but they struggle with long contexts and multi-step instructions. The Minion protocol addresses these weaknesses by decomposing tasks into smaller subtasks, executing them in parallel on device, and aggregating outputs from the cloud. This approach delivers 97.9% of remote-only solution accuracy at a cost of just 17.5%. By leveraging hardware utilization, sequential communication, and model choice, Minions enables a cost-effective and efficient way to distribute AI workloads between small devices and cloud APIs.
Feb 25, 2025
1,257 words in the original blog post.
Together AI has announced a $305 million Series B funding round, led by General Catalyst and co-led by Prosperity7, to scale its AI acceleration cloud for open source and enterprise AI. The investment will accelerate the company's leadership as the preferred AI Cloud for building modern AI applications with open source models, and for training custom models with NVIDIA Blackwell GPUs. This funding will enable Together AI to expand its infrastructure, including securing 200 MW of power capacity and deploying optimized clusters of NVIDIA Blackwell GPUs across multiple North American data centers. The company's research lab continues to pioneer breakthrough methods at the intersection of AI and systems optimization, with innovations like Mixture of Agents, Medusa, Sequoia, Hyena, and Mamba that optimize AI accuracy, performance, and efficiencies. With this investment, Together AI aims to make open source AI accessible to developers and enterprises globally, advancing the frontier of AI through open collaboration, innovation, and transparency.
Feb 20, 2025
808 words in the original blog post.
Together AI has achieved a 90% faster BF16 training with NVIDIA Blackwell Platform and Together Kernel Collection. The team used advanced features like 5th-generation Tensor Cores, on-chip Tensor Memory, and peer CTA groups to develop custom FP8 kernels that run 1.8x faster than FlashAttention-3. This collaboration combines Together AI's kernel optimization expertise with NVIDIA's latest accelerated computing platform innovations, setting new benchmarks for AI training and inference efficiency. The company is deploying tens of thousands of NVIDIA HGX B200 servers and GB200 NVL72 rack-scale solutions to build and deploy the next generation of AI reasoning models and agents. To celebrate NVIDIA Blackwell's arrival, Together AI is offering an exclusive launch program that invites AI teams to apply for a free accelerated test drive of Together GPU Clusters powered by NVIDIA HGX B200 and NVIDIA GB200 NVL72.
Feb 13, 2025
1,422 words in the original blog post.
Together AI is expanding its infrastructure to support large-scale DeepSeek-R1 workloads with Together Reasoning Clusters, which provide dedicated GPU infrastructure for high-throughput, low-latency inference. This offering is designed for companies running large-scale reasoning models and provides benefits such as consistent, low-latency performance, cost-effective scaling, secure environments, and enterprise support. The company also offers a fast, secure serverless API for DeepSeek-R1, with features like instant scalability, flexible pricing, and higher rate limits compared to other providers. Additionally, Together AI provides a free endpoint for the 70B distilled model, allowing users to get started with no upfront cost.
Feb 12, 2025
984 words in the original blog post.