Home / Companies / Together AI / Blog / September 2025

September 2025 Summaries

4 posts from Together AI

Filter
Month: Year:
Post Summaries Back to Blog
The improved Batch Inference API offers significant enhancements, including a streamlined user interface, expanded support for all serverless models and private deployments, and a substantial increase in rate limits from 10 million to 30 billion enqueued tokens per model per user, representing a 3000× increase. This makes it simpler, faster, and more economical, operating at 50% of the cost of the real-time API for processing large-scale datasets. It is particularly advantageous for high-throughput tasks like large-scale text analysis, fraud detection, synthetic data generation, and content moderation, enabling teams like Inception Labs to conduct massive experiments efficiently. These updates aim to make large-scale inference more accessible and cost-effective, positioning the Batch Inference API as an ideal solution for handling extensive workloads without real-time constraints.
Sep 15, 2025 374 words in the original blog post.
Together AI's Fine-Tuning Platform enhances AI developers' ability to customize large language models (LLMs) by offering tools that streamline the training process, allowing for the fine-tuning of models on domain-specific data to improve task performance while reducing costs and latency. The platform supports a range of large models, including those with over 100 billion parameters, and facilitates handling long contexts, crucial for tasks such as long-document processing. Integrations with the Hugging Face Hub enable developers to fine-tune existing models or upload their own, fostering a seamless workflow for model training and deployment. The platform also introduces advanced training objectives for preference optimization and offers convenience features like automatically setting the batch size to maximize efficiency. Together AI's advancements aim to make sophisticated model training more accessible and cost-effective, encouraging developers to integrate fine-tuning into their AI development cycle.
Sep 10, 2025 1,410 words in the original blog post.
AI native startups are rapidly gaining traction, with an expected 378 million users by the end of 2025, but they face challenges in delivering reliable experiences as they scale. Together AI offers a fully managed GPU cloud to help these companies focus on core business rather than managing infrastructure, allowing them to scale seamlessly and efficiently. Trusted by over 800,000 AI engineers, Together AI supports companies like Hedra, Vercept, and SCB 10X with high-performance GPU data centers optimized for AI workloads. Mahadev Konar, a former VP of Infrastructure at Instacart and one of the original architects of Apache Hadoop, has joined Together AI as SVP of Infrastructure Engineering to enhance their GPU infrastructure platform, ensuring reliability, performance, and scalability for AI applications.
Sep 10, 2025 479 words in the original blog post.
Together Instant Clusters are now generally available, providing an API-first, self-service platform for AI infrastructure that automates the setup and management of GPU clusters, ranging from single-node to large multi-node configurations with hundreds of interconnected GPUs. This service is designed to streamline AI workflows by enabling rapid provisioning and scaling without the need for lengthy procurement processes or manual approvals. The clusters come pre-configured with essential components such as NVIDIA drivers, CUDA versions, and networking operators, ensuring they are production-ready and optimized for low-latency inference and high-throughput distributed training. Ideal for AI Native companies facing variable demand, Together Instant Clusters allow for quick scaling to accommodate intense training runs and inference traffic. They offer robust reliability measures to ensure cluster stability and performance, with continuous monitoring and real-time anomaly detection. Pricing is straightforward, with options for hourly, daily, or longer-term usage, and includes free data transfer and reasonably priced shared storage. This innovation aims to enhance productivity and research velocity for AI teams by facilitating faster deployment and operation of high-performance GPU clusters.
Sep 09, 2025 1,061 words in the original blog post.