Home / Companies / Deepinfra / Blog / Post Details
Content Deep Dive

Qwen3.5 122B A10B API Benchmarks: Latency, Throughput & Cost

Blog post from Deepinfra

Post Details
Company
Date Published
Author
Deep
Word Count
1,361
Company Posts That Month
34
Language
English
Hacker News Points
-
Post removed?
No
Summary

Qwen3.5 122B A10B is a sophisticated multimodal foundation model from Alibaba Cloud, designed for applications involving text, image, and video inputs, and featuring a hybrid architecture that utilizes Gated Delta Networks and a sparse Mixture-of-Experts for efficient processing. Released in February 2026, it competes strongly in benchmarks with a remarkable 122 billion parameters, supporting a wide linguistic range of 201 languages and dialects. The model is available through various inference providers, with DeepInfra (FP8) emerging as the top choice due to its superior performance in speed, latency, and cost, although it lacks JSON mode support. DeepInfra offers the lowest blended price at $0.94 per million tokens, the fastest output speed, and the lowest latency, making it the preferred option for most production-scale deployments. However, Alibaba Cloud is recommended for users needing JSON mode, while Novita and GMI provide viable alternatives for specific integration needs, albeit at a higher cost and lower performance compared to DeepInfra.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Agents 1 4,430 1,100 236 -3%
Loop engineering 1 53 37 25 +18%
RAG 1 941 216 85 -48%
Real-time 1 6,296 1,346 246 -2%
Reinforcement learning 1 104 49 23 -14%
Vector Search 1 1,739 413 146 -27%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.