Home / Companies / Together AI / Blog / April 2025

April 2025 Summaries

9 posts from Together AI

Filter
Month: Year:
Post Summaries Back to Blog
At NVIDIA GTC, Together AI announced its collaboration with NVIDIA to accelerate AI workloads on the Blackwell platform. Salesforce leverages Together AI for the entire AI journey, including training, fine-tuning, and inference of their models, and saw a 2x improvement in training speeds when upgrading from HGX H200 to HGX B200. Zoom experienced a 1.9X improvement in training speeds over previous generation NVIDIA Hopper GPUs. InVideo immediately saw a 25% improvement when running a training job on the NVIDIA HGX B200, and then doubled this improvement with optimizations made in partnership with Together AI researchers. The Together Training Stack includes custom-built containers for debugging and running diagnostics at scale, comprehensive MFU benchmarks, full bandwidth benchmarking toolkits, and collective communication diagnostics toolkits. Price-performance is widely considered the most important metric when it comes to GPU cloud infrastructure, and Together AI specializes in delivering higher tokens/sec/node and overall MFU than other providers on the same hardware. The platform offers strong price-performance, infrastructure and security features, technical expertise and support, and flexible consumption models.
Apr 24, 2025 1,145 words in the original blog post.
Chipmunk, a novel training-free method, accelerates diffusion transformers with hardware-aware dynamic column-sparse deltas. By caching attention weights and MLP activations from previous steps, Chipmunk dynamically computes a sparse "delta" against the cached weights. This approach achieves significant speedups in video generation and image generations on various datasets, including up to 3.7x faster video generation at 720x1280 resolution for a 5s video. The method exploits the slow-changing nature of diffusion transformer activations and their inherent sparsity to reduce compute costs. Chipmunk also leverages hardware-efficient sparsity patterns, optimized kernels, and fast cache writeback mechanisms to achieve its performance gains. The technique is designed to be open-sourced and integrated with various model architectures for further acceleration.
Apr 21, 2025 1,550 words in the original blog post.
Together AI has introduced a new platform for fine-tuning language models, enabling businesses to easily refine and improve their models based on user preferences and fresh data. The platform supports preference optimization and continued training, allowing developers to fine-tune open-weight models like Llama or Gemma to reflect users' expectations and capture domain specifics. With the introduction of a new web UI, developers can now start fine-tuning runs directly from the browser, making it more accessible for those without extensive technical knowledge. The platform also includes Direct Preference Optimization, which trains language models on preference data that does not involve an additional reward model. Additionally, the platform supports continued training, allowing developers to update their existing models with new data and adapt them as their app evolves. Other improvements include support for training top-ranking open models, message weights, and learning rate schedulers, as well as optimized data preprocessing logic. The pricing has also been updated to make it more transparent and lower, with no minimum price for fine-tuning. Overall, the platform aims to empower developers and businesses to continuously evolve their models with full ownership.
Apr 17, 2025 1,360 words in the original blog post.
The Together Fine-Tuning Platform now supports Direct Preference Optimization (DPO), a technique to align language models with human preferences, creating more helpful and accurate AI assistants. DPO allows training directly on preference data without an intermediate reward model, simplifying the process compared to traditional approaches like Reinforcement Learning from Human Feedback (RLHF). This method refines how capabilities are expressed in the model, improving its generation quality and alignment with human values. DPO is ideal for tasks where prompting isn't sufficient or when humans can compare better than create, making controlled improvements to existing models more efficient. The technique excels in tasks with nuanced quality judgments, but may not be suitable for single correct answers or tasks with objectively correct answers. To get started with DPO on Together, developers need to tune key hyperparameters like --dpo-beta and monitor training metrics specific to preference optimization.
Apr 17, 2025 1,472 words in the original blog post.
Continued fine-tuning of large language models allows for sequential fine-tuning by specifying the --from-checkpoint parameter, enabling builds upon previously trained models. This process is crucial in adapting models to new tasks, domains, or languages while preserving their existing capabilities. Continued fine-tuning offers a resource-efficient way to adapt models to changing requirements without sacrificing previous learned skills. It encompasses various approaches, including fine-tuning for different tasks, instruction tuning, model refinement, and alignment. The key challenge is catastrophic forgetting, which can be mitigated by using similar task datasets across languages. Continued fine-tuning is valuable in scenarios where a model needs to adapt to new data or tasks, incorporate new knowledge incrementally, or align with human preferences. It requires careful consideration of dataset similarity, learning rate, and performance metrics. The approach has promising applications in enhancing multilingual capabilities and improving task-specific performance.
Apr 17, 2025 1,292 words in the original blog post.
The Open Deep Research workflow is an emerging AI technology that enhances web search by producing comprehensive, well-cited content on complex topics. It features a flexible architecture designed to educate developers and be extended by the community, focusing specifically on providing in-depth understanding of specific topics rather than just short answers to complex questions. The workflow integrates models from the TogetherAI cloud platform, with each component - text, image, and audio generation - powered by the most suitable model available through TogetherAI's cloud infrastructure. The implementation aims to improve upon current approaches and share practical insights into what works and what doesn't when building such complex AI workflows from scratch.
Apr 16, 2025 3,100 words in the original blog post.
DeepCoder-14B-Preview, a fully open-source 14B coder, achieves an impressive 60.6% Pass@1 accuracy on LiveCodeBench, matching the performance of o3-mini-2025-01-031 and o1-2024-12-17 with just 14B parameters. The model was trained using a curated high-quality training set consisting of TACO Verified problems, PrimeIntellect's SYNTHETIC-1 dataset, and LiveCodeBench problems submitted between May 1, 2023, and July 31, 2024. To accelerate end-to-end RL training, the authors introduce verl-pipeline, an optimized extension of the open-source RLHF library Verl, which achieves up to 2.5× speedup over the baseline implementation. The model demonstrates strong performance across various coding benchmarks, including LiveCodeBench, Codeforces, and HumanEval+, achieving 60.6% on LiveCodeBench and a rating of 1936 on Codeforces, comparable to the performance of o3-mini (low) and o1.
Apr 08, 2025 2,870 words in the original blog post.
Together AI has partnered with Meta to offer Llama 4, a cutting-edge mixture-of-experts (MoE) architecture model that combines native multimodality and offers unparalleled efficiency and scale. Two models are available: Llama 4 Maverick, which excels at multilingual image/text understanding, creative writing, and enterprise-scale applications, and Llama 4 Scout, which is optimized for multi-document analysis, codebase reasoning, and personalized tasks. The models offer improved performance compared to previous versions, with lower costs, faster training times, and higher network compression rates. Developers can deploy the models on Together AI's platform or use them directly through their API, allowing for greater control over their data and models.
Apr 05, 2025 608 words in the original blog post.
Dippy AI, a startup founded in April 2024, has grown to over 4 million users creating and publishing unique AI characters and exchanging messages with them, but faced challenges managing infrastructure at scale. To address this, they partnered with Together AI engineers to deploy their custom models on Together Dedicated Endpoints, leveraging optimized GPU infrastructure to handle volumes of 4M+ tokens/minute with optimal "throughput per dollar" without spending time managing the infrastructure. This partnership enabled Dippy AI's team to focus on building user-facing features and enhancing the user experience, while reducing latency and cost. With the out-of-the-box auto-scaling of Together Dedicated Endpoints, Dippy experienced predictable, steady availability, with no capacity issues, resulting in consistent, uninterrupted interactions for their users. The partnership has allowed Dippy AI to meet and improve KPIs such as Time to First Token, Throughput, and Latency, while enabling the team to focus on improving the product and user experience. Together Dedicated Endpoints have also provided lower cost, faster training, and network compression, supporting upcoming innovations like voice calls and state-of-the-art AI audio models.
Apr 01, 2025 1,074 words in the original blog post.