Home / Companies / Baseten / Blog / Post Details
Content Deep Dive

How we built the fastest GLM 5 API

Blog post from Baseten

Post Details
Company
Date Published
Author
Tri Dao 2 others
Word Count
858
Company Posts That Month
9
Language
English
Hacker News Points
-
Post removed?
No
Summary

Z.ai's GLM-5, an open-weight model developed by Andon Labs, has achieved state-of-the-art results in both time to first token (TTFT) and tokens per second (TPS) with its innovative use of a mixture of experts (MoE) architecture, which selectively activates parameters based on the task at hand. This model, which is more than twice the size of its predecessor GLM-4.7, excels in tasks such as code generation and agentic reasoning, and ranks highest among open-source models in the Vending Bench 2 benchmark, which assesses a model's decision-making capabilities over a long time horizon. By leveraging custom kernels optimized for DeepSeek Sparse Attention and a low-overhead Multi-Token Prediction (MTP) speculative decoding engine, GLM-5 achieves 186+ tokens per second, making it the fastest in inference as benchmarked by Artificial Analysis. The Baseten Inference Stack enhances performance through KV-aware routing, MoE dispatch kernel optimizations, and NVFP4 quantization for compatibility with Blackwell inference. These innovations underscore GLM-5's suitability for complex systems engineering and autonomous coding tasks, offering industry-leading throughput for open-source models.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 1 6,078 960 218 +18%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.