October 2025 Summaries
9 posts from Baseten
Filter
Month:
Year:
Post Summaries
Back to Blog
Baseten Training, now generally available, addresses key challenges in AI model training by providing a flexible, developer-friendly infrastructure that eliminates the complexities of managing compute resources, pricing, and operational control. This platform allows developers to train any AI model on any dataset effortlessly, offering tools like caching, checkpointing, and deployment options that enhance performance and reduce costs, demonstrated by OpenEvidence's reported 23x speed improvement and significant cost savings. The platform supports a wide range of training jobs, from small finetunes to large-scale reinforcement learning projects, and integrates seamlessly into existing workflows, allowing developers to focus on model training rather than infrastructure management. Baseten Training's enhancements since its beta launch include a collection of open-source recipes, improved UI, and more robust logging and observability features, enabling users to efficiently manage training projects and deploy models with ease.
Oct 30, 2025
922 words in the original blog post.
NVIDIA's Nemotron Nano 2 VL is an advanced open-source vision language model designed for high performance and scalability, particularly in the financial services sector, where it excels in tasks such as know-your-customer compliance, intelligent document processing, and fraud detection. Built on a base of the Nemotron Nano 2, a 9 billion parameter foundation model, Nemotron Nano 2 VL features a 12 billion parameter architecture with a hybrid Mamba-Transformer design, offering enhanced accuracy and efficiency for tasks like multi-image understanding, document intelligence, and video captioning. This model, available on Baseten, supports robust inference capabilities using NVIDIA NIM microservices for high throughput and low latency performance, with applications extending to various industries such as healthcare and media. Additionally, a smaller model, Nemotron Parse 1.1, provides a cost-effective solution for straightforward optical character recognition tasks. Baseten enhances the deployment of these models with enterprise-grade security, multi-cloud infrastructure, and technical support, making them suitable for building secure, reliable AI agents in production environments.
Oct 28, 2025
871 words in the original blog post.
Baseten's advancements in speculative decoding, particularly with EAGLE-3, have significantly increased the performance of their GPT-OSS 120B inference API, achieving over 650 tokens per second, as verified by Artificial Analysis. As a launch partner for GPT-OSS, Baseten has utilized powerful NVIDIA hardware, including B200 GPUs, TensorRT-LLM, and NVIDIA Dynamo, to establish itself as the fastest NVIDIA-based provider, rivaling custom hardware providers. The implementation of techniques such as tensor parallelism and KV-aware routing have further optimized performance, while the potential for additional improvements through PD disaggregation and advanced speculation remains under investigation. Benchmarks from Artificial Analysis and OpenRouter highlight Baseten's superior performance, showcasing their ability to match or exceed custom hardware solutions while maintaining flexibility and scalability with NVIDIA GPUs. Baseten continues to innovate in model performance engineering and offers opportunities for further exploration and employment in this field.
Oct 24, 2025
1,188 words in the original blog post.
DeepSeek-OCR is an innovative Optical Character Recognition model that revolutionizes data processing by utilizing a unique compression technique, reducing the need for visual tokens by tenfold compared to traditional text tokens, with a decoding precision of 97%. This efficiency not only allows the model to process vast amounts of data quickly and cost-effectively but also suggests a broader impact on AI intelligence by improving data representation for downstream tasks. The model's implementation on Baseten, using Truss and vLLM, demonstrates its scalability and reliability, even when faced with challenging inputs like doctors' handwriting. This approach highlights a shift in AI data processing from text to visual tokens, underscoring the potential for developing real-time AI agents and advancing document retrieval and question answering systems. The deployment process on Baseten, involving specific configurations and dependencies, illustrates the ease of integrating DeepSeek-OCR for various applications, offering a pathway to harness its capabilities for scalable training data generation and enhancing AI contextual understanding.
Oct 24, 2025
988 words in the original blog post.
Baseten collaborates with NVIDIA to enhance model performance through the adoption of NVIDIA Dynamo, an open-source inference framework designed for large-scale LLM serving across distributed GPU clusters. A key feature of NVIDIA Dynamo is its KV cache-aware routing, which optimizes inference speeds by directing requests to model replicas with cached contexts, significantly reducing redundant computations and improving system performance. This routing approach, which balances cache hit rates and workload distribution, leads to substantial improvements in metrics such as time to first token (TTFT) and time per output token (TPOT), as demonstrated in benchmarks with models like Qwen3 Coder. Baseten has observed notable reductions in latency and increases in throughput, processing more requests per second and outputs per second. Looking ahead, Baseten plans to further leverage NVIDIA's tools, exploring features like KV cache offloading to enhance resource utilization and concurrency, and they will co-host a technical workshop with NVIDIA to share insights on maximizing AI inference workloads using NVIDIA Dynamo.
Oct 17, 2025
904 words in the original blog post.
Autodesk's WaLa model is a groundbreaking single-view 3D reconstruction technology that transforms 2D images into detailed 3D models using wavelet-based latent diffusion, capable of producing high-quality OBJ meshes suitable for 3D printing or integration into game engines. The text outlines a tutorial on transforming this research model into a production API using Truss and Netlify, creating a scalable system that converts sketches into shareable 3D flower cards. It emphasizes the process of deploying the model as a scalable API on Baseten, which includes containerization and dependency management via Truss, and storing and viewing 3D models through Netlify's serverless functions and database integration. The resulting application offers an interactive web interface for drawing and transforming sketches into 3D models, complete with shareable URLs and a mobile-friendly, touch-responsive viewer built with Three.js. The project demonstrates the potential of turning open-source models into practical applications, with all code made available on GitHub for further experimentation and development.
Oct 11, 2025
1,457 words in the original blog post.
In an interview with Dax Raad, creator of OpenCode and Zen, various topics are explored, including the motivation behind developing OpenCode, the open-source philosophy, and the importance of terminal-based workflows. OpenCode was developed to address the cumbersome experience of using large language models (LLMs) in traditional environments, by integrating LLMs directly into the terminal where users can interact with their file systems seamlessly. Dax explains the necessity of open-source for OpenCode to leverage community contributions and adapt to the constantly evolving landscape of LLMs, contrasting this with closed-source solutions like Claude Code that do not benefit as much from community input. Zen is introduced as a complementary service to OpenCode, providing access to high-quality deployments of popular and new models in an economically sustainable way. Dax also critiques the over-reliance on benchmarks for evaluating AI products, arguing that real-world usability and user experience are more significant indicators of product quality, and highlights the "superstitious behavior" people develop around AI models due to their unpredictable nature.
Oct 10, 2025
2,827 words in the original blog post.
Baseten's text-to-video system, running on Nebius using the Baseten Inference Stack, is designed to deliver predictable and efficient performance by integrating advanced technologies and infrastructure. The system's Inference Runtime utilizes custom modality-specific kernels, kernel fusion, and attention kernels to optimize video workloads, along with topology-aware parallelism and continuous batching to manage request prioritization and latency. Inference-optimized Infrastructure ensures reliable and scalable performance through intelligent request routing, geo-aware load balancing, SLA-aware autoscaling, and active-active reliability. These components work together to maintain consistent latency and throughput across varying traffic levels, enabling efficient scaling without compromising quality. The use of Nebius's large GPU pools and low-friction capacity growth complements Baseten's sophisticated runtime and infrastructure, ensuring a seamless experience even during demand spikes. The system's ability to adapt to real-time changes and manage multi-cloud capacity efficiently makes it robust enough to handle complex video generation tasks, turning a demo into a fully-fledged product offering.
Oct 06, 2025
867 words in the original blog post.
Healthcare organizations are increasingly integrating AI into their operations to reduce costs, improve diagnostics, and enhance patient experiences, necessitating infrastructure that supports low latency, robust security, and scalability. AI teams are focusing on deploying and refining open-source models or developing custom models tailored to specific healthcare tasks, such as document processing, clinical assistance, and diagnostic image recognition, which require high availability and cost-effectiveness. To meet these demands, Baseten and Vultr provide a secure, HIPAA-compliant infrastructure using NVIDIA HGX B200 systems, enabling efficient AI model inference with flexible access to GPUs. This collaboration supports healthcare AI engineering teams in overcoming production challenges and achieving rapid market deployment with controlled costs, ensuring that AI solutions can handle high traffic volumes while maintaining performance and security.
Oct 02, 2025
823 words in the original blog post.