October 2026 Summaries
5 posts from Baseten
Filter
Month:
Year:
Post Summaries
Back to Blog
No summary generated yet.
Oct 09, 2026
792 words in the original blog post.
No summary generated yet.
Oct 09, 2026
1,759 words in the original blog post.
No summary generated yet.
Oct 09, 2026
879 words in the original blog post.
Baseten reports that its inference stack achieved more than 200 tokens per second and roughly 200 ms time to first token for Z.ai’s open-weight GLM-5 model in Artificial Analysis benchmarks. GLM-5 is a 744-billion-parameter mixture-of-experts reasoning model with 40 billion active parameters per pass, a 200K-token context window, MIT licensing, and strong reported coding, mathematics, and agentic-task benchmark results. Baseten attributes its performance to architecture-specific MoE routing and dispatch kernels for GLM-5’s deeper, narrower design, custom support for DeepSeek Sparse Attention to reduce long-context KV-cache costs, and low-overhead speculative decoding using the model’s native Multi-Token Prediction heads rather than a separate draft model. The company argues that throughput is especially important for reasoning workloads, where hidden thinking sequences may be substantially longer than final answers, and positions the optimized deployment for coding agents, tool-use pipelines, and other long-running production applications.
Oct 02, 2026
1,009 words in the original blog post.
Baseten reports testing MetaInfer, a skills-based framework that uses LLM agents to build specialized inference engines without model post-training, finding that highly constrained engines can exceed general-purpose serving systems for particular model, hardware, and workload combinations. Using Claude Code with Fable 5, access to a B200 GPU, reference model weights for accuracy checks, and production-style AIPerf evaluations, the team generated VibeQwen for Qwen-3.6-35B-A3B in NVFP4 precision, which reportedly achieved 90% faster single-stream decoding than tuned vLLM 0.25.1, reduced time to first token from 28 ms to 12 ms, and delivered 71% greater output throughput at concurrency 32. A second experiment reused the expanded knowledge base to create Sammie, a server for the SAM 3.1 image-segmentation model that processed 91 images per second on an H100, a reported 50% improvement over Meta’s reference server. The author argues that autonomous optimization is increasingly feasible because inference performance and correctness can be quantitatively measured, although carefully designed accuracy and workload constraints are necessary to prevent agents from optimizing misleading metrics. Both systems remain experimental rather than production-serving, but the reported results suggest that AI-assisted, reusable optimization workflows could become a routine component of model deployment as model capabilities improve and costs decline.
Oct 02, 2026
1,898 words in the original blog post.