Home / Companies / Together AI / Blog / July 2025

July 2025 Summaries

9 posts from Together AI

Filter
Month: Year:
Post Summaries Back to Blog
VirtueGuard, developed by Virtue AI, is now available on Together AI, offering a pioneering solution in AI security and safety that is both comprehensive and efficient for real-world applications. This guardrail model boasts an impressive 8ms response time, performing 50 times faster than alternatives without compromising accuracy, while ensuring 99.9% uptime for enterprise reliability. Created by renowned AI security experts, VirtueGuard addresses the crucial need for organizations to deploy AI models without risking harmful outputs, compliance violations, or reputational damage. It seamlessly integrates into Together AI's platform, providing protection across text, images, and audio by monitoring 12 risk categories, including privacy violations and hate speech, with a simple API parameter. Already trusted by leading companies such as Uber and NVIDIA, VirtueGuard exemplifies cutting-edge AI security with continuous updates to threat models and policy frameworks, ensuring enterprises receive automatic, up-to-date protection.
Jul 29, 2025 608 words in the original blog post.
In the rapidly evolving field of large language models (LLMs), assessing a model's performance on specific tasks is essential, and Together Evaluations provides a structured way to benchmark these models using LLMs as judges. This approach allows for fast and flexible evaluation by defining task-specific benchmarks and employing leading open-source models to compare responses, bypassing the need for manual labeling or rigid metrics. Together Evaluations supports three modes of evaluation—classify, score, and compare—each customizable via prompt templates, allowing users to tailor the process to their specific needs. This system enables developers to identify the best models or prompts, monitor model quality, and manage data drift, with support for serverless inference and the ability to upload pre-existing datasets for evaluation. By offering comprehensive tools and demonstrations, Together aims to streamline the development of LLM-driven applications and invites users to explore the platform through interactive resources and a webinar.
Jul 28, 2025 1,176 words in the original blog post.
Qwen3-Coder, available on Together AI's platform, is an advanced coding model designed to handle complex and interconnected software engineering tasks, boasting 480 billion parameters and the ability to handle entire codebases rather than isolated snippets. It delivers frontier-level performance in real engineering workflows, such as legacy system modernization, cross-system feature development, and complex debugging, outperforming traditional models in these areas. Together AI's infrastructure is specifically optimized for AI workloads, offering performance, reliability, and security without the trade-offs seen in other cloud services, and allows for instant deployment with zero setup. The model's capabilities continue to improve over time thanks to continuous optimizations, making it ideal for development teams needing to modernize authentication systems, refactor architectures, or implement complex features across multiple services.
Jul 25, 2025 643 words in the original blog post.
FutureBench is a proposed benchmarking framework aimed at evaluating artificial intelligence models based on their ability to predict future events, rather than just relying on past information or static datasets. This approach emphasizes the importance of sophisticated reasoning, synthesis, and genuine understanding, as opposed to mere pattern matching, by using real-world prediction markets and live news to generate meaningful prediction tasks across various domains like science, economics, and geopolitics. By focusing on forecasting, FutureBench addresses challenges of data contamination common in traditional benchmarks and creates a more objective, verifiable measure of model performance. The framework operates on three levels—comparing agentic frameworks, tool performance, and model capabilities—allowing for a comprehensive analysis of how models gather and synthesize information to make predictions. Initial results have demonstrated varying strategies and reasoning patterns among different models, revealing insights into their information-gathering behaviors and decision-making processes. The benchmark is seen as a dynamic tool, evolving with community feedback to refine its question sourcing and experimental methods, although it faces challenges like high costs due to the extensive use of input tokens.
Jul 17, 2025 1,867 words in the original blog post.
Together AI has introduced support for NVIDIA Blackwell GPUs in their inference platform, enhancing the performance of AI models like DeepSeek-R1-0528, particularly when deployed on NVIDIA HGX B200 GPUs. This advancement positions Together AI as a leader in high-speed AI inference, leveraging a combination of bespoke GPU kernels, a proprietary inference engine, and innovative techniques such as speculative decoding and lossless quantization. The platform demonstrates notable speed improvements, achieving up to 334 tokens per second, and offers customizable Dedicated Endpoints for further optimization in production environments. The improved performance is achieved without compromising model quality, thanks to Together AI's advanced inference stack that includes state-of-the-art components and methodologies. This development enables efficient and scalable deployment of AI workloads, with Together AI providing options for both serverless and dedicated endpoints to meet diverse customer needs.
Jul 17, 2025 1,527 words in the original blog post.
Kimi K2, available on Together AI's cloud platform, is a leading open-source AI model with 1 trillion parameters, offering exceptional performance across a wide range of applications. Its notable achievements include ranking #1 in creative writing and autonomous reasoning, and it excels in tool mastery and production-ready coding. Kimi K2 is designed for agentic use, seamlessly executing complex workflows by analyzing data, generating insights, and creating reports without requiring detailed instructions. Its architecture utilizes a Mixture-of-Experts design and has been trained on 15.5 trillion tokens. This model is deployed serverlessly, ensuring high reliability, instant scalability, and cost-effectiveness, making it competitive with proprietary models while offering superior economics. The platform supports diverse applications, from customer support and content generation to research and model development, offering a secure and scalable environment for enterprise deployment.
Jul 14, 2025 910 words in the original blog post.
Together AI has launched its new speech-to-text APIs, addressing the speed and quality challenges faced by voice application developers. Their Whisper V3 Large deployment offers transcription 15 times faster than OpenAI while maintaining accuracy, thanks to optimizations like smart voice activity detection, intelligent chunking, and improved GPU utilization. These advancements enable real-time applications in sectors like customer support, meetings, and healthcare, by eliminating the traditional bottlenecks associated with audio processing. The APIs support files over 1GB, provide superior word-level alignment, and handle over 50 languages, offering substantial cost savings for high-volume applications. The service is designed for easy integration, with compatibility for existing Whisper users and an interactive playground for real-time testing. This release marks a significant step towards building a comprehensive voice infrastructure, making voice-enabled applications faster and more accessible.
Jul 10, 2025 592 words in the original blog post.
Together AI has achieved SOC 2 Type 2 compliance, demonstrating its commitment to security, privacy, and regulatory standards, which bolsters confidence in deploying advanced AI workloads. The rigorous compliance process included an independent audit that verified the effectiveness of security protocols such as access management, data encryption, and incident response. The company employs a layered security architecture with features like network segmentation and automated threat detection, supported by multi-factor authentication and role-based access controls. Together AI's adherence to HIPAA requirements further allows healthcare entities to utilize its AI platform for applications like clinical decision support and biomedical data analysis. The organization emphasizes continuous security improvements through regular assessments and an incident response plan, showcasing its dedication to technical excellence and regulatory adherence, which supports a wider range of use cases in regulated industries.
Jul 08, 2025 325 words in the original blog post.
DeepSWE-Preview, a state-of-the-art coding agent developed through a collaboration between the Agentica team and Together AI, achieves significant performance in reasoning-enabled coding tasks using only reinforcement learning (RL) on the Qwen3-32B model. This open-source agent demonstrates a 59% success rate on the SWE-Bench-Verified benchmark, surpassing previous open-weight models with 42.2% Pass@1 and 71.0% Pass@16 scores. The agent is trained through Agentica's rLLM framework, utilizing 4,500 real-world software engineering tasks over six days on 64 H100 GPUs, and the entire process, including datasets, code, and training logs, is open-sourced for community advancement. DeepSWE-Preview innovatively navigates complex software engineering environments, leveraging a mix of reinforcement learning techniques and hybrid test-time scaling strategies to enhance coding agents' efficacy. The project underscores the potential of RL to advance long-horizon, multi-step reasoning models in software development, offering a comprehensive foundation for future explorations in agentic AI domains.
Jul 02, 2025 3,655 words in the original blog post.