Home / Companies / Deepinfra / Blog / Post Details
Content Deep Dive

GLM-4.7-Flash API Benchmarks: Latency, Throughput & Cost

Blog post from Deepinfra

Post Details
Company
Date Published
Author
Deep
Word Count
1,455
Company Posts That Month
34
Language
English
Hacker News Points
-
Post removed?
No
Summary

GLM-4.7-Flash, developed by Z.AI and released in January 2026, is an open-source reasoning model based on a Mixture-of-Experts Transformer architecture with 30 billion parameters, designed for efficient performance in agentic workflows and multi-step reasoning tasks. This model demonstrates state-of-the-art performance among open-source models in its size category, supporting up to 200K context tokens and enabling deployment on consumer hardware. The analysis of various inference providers reveals that DeepInfra offers the best overall value for GLM-4.7-Flash deployment, providing the lowest latency at 0.75 seconds, the cheapest cost at $0.14 per million tokens, and full support for JSON Mode and Function Calling, making it particularly suitable for real-time applications. Amazon Bedrock is noted for its superior throughput, making it ideal for high-volume batch processing despite its lack of JSON Mode support. In contrast, Novita is not recommended for production use due to high latency issues, although it shares feature support with DeepInfra.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 2 6,296 1,346 246 -2%
AI Agents 1 4,430 1,100 236 -3%
LLM 1 5,932 1,046 223 -2%
Vector Search 1 1,739 413 146 -27%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.