How we built the fastest GLM 5 API
Blog post from Baseten
Z.ai's GLM-5, an open-weight model developed by Andon Labs, has achieved state-of-the-art results in both time to first token (TTFT) and tokens per second (TPS) with its innovative use of a mixture of experts (MoE) architecture, which selectively activates parameters based on the task at hand. This model, which is more than twice the size of its predecessor GLM-4.7, excels in tasks such as code generation and agentic reasoning, and ranks highest among open-source models in the Vending Bench 2 benchmark, which assesses a model's decision-making capabilities over a long time horizon. By leveraging custom kernels optimized for DeepSeek Sparse Attention and a low-overhead Multi-Token Prediction (MTP) speculative decoding engine, GLM-5 achieves 186+ tokens per second, making it the fastest in inference as benchmarked by Artificial Analysis. The Baseten Inference Stack enhances performance through KV-aware routing, MoE dispatch kernel optimizations, and NVFP4 quantization for compatibility with Blackwell inference. These innovations underscore GLM-5's suitability for complex systems engineering and autonomous coding tasks, offering industry-leading throughput for open-source models.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 1 | 6,078 | 960 | 218 | +18% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.