March 2025 Summaries
8 posts from Together AI
Filter
Month:
Year:
Post Summaries
Back to Blog
Together AI has been awarded a ClusterMAX Gold rating by SemiAnalysis, a prestigious evaluation of GPU cloud providers. This achievement highlights Together AI's rapid innovation, exceptional support, and deep industry expertise in building and managing cutting-edge GPU infrastructure. The award is based on strong security practices, reliable infrastructure, competitive pricing, and deep technical support, including expertise in optimizing GPUs with proprietary kernels, such as the Together Kernel Collection (TKC). The ClusterMAX Gold rating underscores Together AI's commitment to delivering a flexible ecosystem for developers, with features like managed Slurm or Kubernetes solutions, and its partnership with NVIDIA to deliver the latest hardware and direct lines to the NVIDIA team for further optimizations. Together AI is also rapidly working to improve its monitoring and alerting capabilities, as well as exploring new technologies like SHARP in-network reduction on its InfiniBand fabric. The award marks a significant milestone in Together AI's trajectory, with the company poised to achieve Platinum status in the next ClusterMAX evaluation.
Mar 27, 2025
973 words in the original blog post.
Together Chat, a next-generation consumer app, has been launched with free access to DeepSeek R1, securely hosted in North America. The app allows users to interact with various open-source AI models, including text-to-image models like Flux Schnell and code generation tools like Qwen Coder 32B. Additionally, Together Chat offers web search capabilities, reasoning models for complex questions, and a mobile app experience. Users can access these features without paying any costs, as the service is completely free to use, with daily credits available for each model. The app's development is built using Together AI APIs, which also enable developers to power LLM features in their own apps.
Mar 24, 2025
648 words in the original blog post.
Frontier AI teams can now deploy high-performance GPU clusters in minutes, thanks to Together Instant GPU Clusters accelerated by up to 64 NVIDIA GPUs per cluster. This self-service solution offers flexible deployment options, including Kubernetes or Slurm for workload orchestration, and allows users to choose their own NVIDIA Driver and CUDA versions. The service also provides a simple and transparent pricing structure, with competitive costs for high-performance compute. With instant provisioning in minutes, users can skip lengthy approvals and procurement cycles, deploy NVIDIA GPUs instantly without waiting for sales conversations or capacity planning, and accelerate AI workloads with ultra-low-latency, high-throughput performance.
Mar 18, 2025
800 words in the original blog post.
This week, as an NVIDIA GTC Gold Sponsor, we're taking the wraps off our latest advancements to Together GPU Clusters. We’re rapidly deploying NVIDIA Blackwell GPUs at massive scale, bringing next-generation AI performance to our customers. Early adopters—including Zoom, Salesforce, and fal.ai—have test driven NVIDIA HGX B200 on Together AI, already seeing an impressive ~2x speed-up in training and inference performance relative to NVIDIA HGX H100—with additional acceleration coming soon. Today, we’re excited to introduce the Preview release of Together Instant GPU Clusters—on-demand clusters of up to 64 NVIDIA GPUs, interconnected via NVIDIA Quantum-2 InfiniBand and NVIDIA NVLink™. Fully self-service, these clusters can be provisioned via our console in just minutes, unlocking powerful AI infrastructure at unprecedented speed and scale. Our customer success is a testament to how our vision for purpose-built and truly open AI infrastructure resonates with both AI-native startups and established enterprise companies, ensuring they have the flexibility and performance needed to push AI forward.
Mar 18, 2025
1,836 words in the original blog post.
Together AI, a leader in scaling generative AI model deployment, has integrated NVIDIA NIM (Neural Inference Model) microservices for secure and reliable high-performance AI model inferencing. This integration enables developers to explore, test, and deploy select NIMs as dedicated endpoints on the Together platform with fast performance and industry-leading cost efficiency. The platform delivers exceptional scale and capacity, a developer-centric experience, cost-effective resource management, and enterprise-ready performance, making it an attractive option for organizations looking to leverage NVIDIA NIM microservices as hosted endpoints.
Mar 18, 2025
744 words in the original blog post.
The ThunderKittens framework, developed in collaboration with Stanford researchers, has been optimized for NVIDIA Blackwell GPUs. The new kernels allow for faster matrix multiplication and attention computations on the B200 architecture, leveraging features such as fifth-generation tensor cores, tensor memory, and CTA pairs. These optimizations enable better dataflow management, reduced bubbles in the pipeline, and increased throughput. By utilizing these new features, developers can write more efficient and performant GPU kernels, making it easier to quickly write high-performance code for various applications.
Mar 15, 2025
1,573 words in the original blog post.
Together AI is announcing the availability of on-demand Dedicated Endpoints, which offer unmatched price-performance for scaling AI inference in production. This new service provides a balance between flexibility and affordability, making it an ideal solution for startups and large companies alike. With up to 43% lower pricing than previous offerings, Dedicated Endpoints deliver high performance, full control and customizability over the deployment hardware and configuration, support for custom fine-tuned models, no minimum commitments, and no upfront costs. The service allows users to spin up on-demand Dedicated Endpoints for popular open-source models or upload their own custom fine-tuned model from Hugging Face, deploy it instantly, and start running inference without any additional storage or upload fees. This update is expected to provide significant cost savings at scale, making it an attractive option for mission-critical AI applications that require reliable QPS, predictable availability, and seamless handling of surges without performance dips.
Mar 13, 2025
1,191 words in the original blog post.
Together AI has partnered with NVIDIA to become an NVIDIA Cloud Partner, offering accelerated AI offerings by leveraging the NVIDIA Blackwell platform. This partnership provides cloud infrastructure designed for fast training and inference performance, reliable scalable service with proactive monitoring, and access to platforms with the latest NVIDIA Cloud Partner reference architecture. Together AI also offers over 200 MW datacenter capacity and is a Gold sponsor at GTC in San Jose, where they will meet with attendees interested in leveraging the NVIDIA Blackwell platform. As an NVIDIA Cloud Partner, Together AI has been optimizing and operating accelerated computing clusters using the NVIDIA Cloud Partner reference architecture, ensuring optimized performance, enterprise-grade reliability, and seamless scalability for AI workloads of any size. They now also offer the complete NVIDIA AI Enterprise software suite to their enterprise customers, providing a comprehensive platform with various frameworks and libraries across many domains. Together AI will have a major presence at GTC 2025, delivering featured presentations on optimizing large language model inference, accelerating AI with purpose-built cloud infrastructure, and revolutionary improvements in attention mechanisms.
Mar 11, 2025
744 words in the original blog post.