May 2025 Summaries
4 posts from Baseten
Filter
Month:
Year:
Post Summaries
Back to Blog
Baseten is a company that powers the next generation of AI-powered products, focusing on providing infrastructure that supports AI growth. As of 2025, Baseten's technology is used by top AI companies to deliver fast, reliable, and cost-effective inference across millions of users. The company has built an extensive inference stack and plans to make exciting product announcements in the near future. Baseten prioritizes details in all aspects of its operations, including customer communication, design, and visual identity, which was recently revamped by a new design system, logo, and style.
May 25, 2025
258 words in the original blog post.
Baseten is introducing two new products, Model APIs and Training, to help customers integrate open-source models into production. The company aims to provide an easy path for developers to use open-source models with state-of-the-art performance, production-grade reliability, and a user-friendly experience. Baseten Model APIs are designed to solve the problems of accessing open-source models in production, while Baseten Training is addressing the challenges of training and fine-tuning AI models. The company has partnered with several companies to offer early access to its new products, which include four great models for launch and plans to add more as the landscape evolves.
May 24, 2025
525 words in the original blog post.
Qwen 3, a new family of open-source LLMs by Alibaba, introduces Qwen 3 235B, a state-of-the-art reasoning model that rivals DeepSeek-R1 but requires significantly fewer hardware resources to run in production. The model uses a Mixture of Experts (MoE) architecture with 128 experts and 8 experts per token, optimized for deployment with SGLang, an open-source fast inference framework. Qwen 3 achieves very usable performance on day zero, with smaller batches reducing latency but increasing the effective cost per token by lowering throughput. The model performs well on public benchmarks, comparing favorably to larger models like DeepSeek-R1 and Gemini 2.5 Pro. To take full advantage of its efficient performance in production, Qwen 3 can be served with low latency and high throughput using SGLang, which automatically splits the model appropriately using the --tp argument to specify the number of GPUs to use for inference. The model's performance can be further improved through various techniques, including varying temperature between thinking modes, taking advantage of agentic capabilities, and using the entire context window. Qwen 3 is now available in both FP8 and BF16 precisions, with recommendations to run in FP8 as it offers nearly identical quality at a much lower cost.
May 19, 2025
1,303 words in the original blog post.
Canopy Labs has selected Baseten as its preferred inference provider for Orpheus TTS models. This partnership enables developers to use the high-performance Orpheus model in production, with optimized performance and scalability on a single H100 MIG GPU. The collaboration between Canopy Labs and Baseten resulted in the creation of the world's highest-performance Orpheus inference server based on NVIDIA's TensorRT-LLM. This allows for 16 concurrent live connections with variable traffic, 24 concurrent live connections with stable traffic, and up to 60x real-time factor for bulk jobs. The client code example provided by Baseten supports session re-use, reducing overhead and improving TTFB performance. With this partnership, developers can now build fast, configurable, and cost-efficient voice agents using Orpheus TTS models.
May 07, 2025
1,350 words in the original blog post.