Home / Companies / Baseten / Blog / Post Details
Content Deep Dive

How we built the new fastest API for GLM-5.2

Blog post from Baseten

Post Details
Company
Date Published
Author
Alex Korte, Magdy Saleh, Tri Dao, Anant Desai, Bryce Dubayah, Abu Qader, Philip Kiely
Word Count
524
Company Posts That Month
12
Language
English
Hacker News Points
-
Post removed?
No
Summary

Baseten recently released a highly optimized API for the GLM-5.2 model, achieving impressive speeds of up to 280 tokens per second and an average of 100 tokens per second, with performance more than doubling since its launch. This includes a fast API version designed for reduced latency, using Tensor and Expert Parallelism, which trades throughput for speed and consequently has higher token prices. The APIs, running on NVIDIA B200 GPUs, have been enhanced through optimizations in scheduling, NVFP4 weights, and speculative decoding profiles, leading to significant reductions in batch size to minimize resource competition. These developments have garnered positive market feedback, confirming the API's superiority in both benchmark tests and real-world application. Further improvements are planned, and the GLM-5.2 Fast API is available for public use on Baseten, with ongoing learnings from the Kimi K3 project expected to inform future enhancements.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.