Qwen3.5 122B A10B API Benchmarks: Latency, Throughput & Cost
Blog post from Deepinfra
Qwen3.5 122B A10B is a sophisticated multimodal foundation model from Alibaba Cloud, designed for applications involving text, image, and video inputs, and featuring a hybrid architecture that utilizes Gated Delta Networks and a sparse Mixture-of-Experts for efficient processing. Released in February 2026, it competes strongly in benchmarks with a remarkable 122 billion parameters, supporting a wide linguistic range of 201 languages and dialects. The model is available through various inference providers, with DeepInfra (FP8) emerging as the top choice due to its superior performance in speed, latency, and cost, although it lacks JSON mode support. DeepInfra offers the lowest blended price at $0.94 per million tokens, the fastest output speed, and the lowest latency, making it the preferred option for most production-scale deployments. However, Alibaba Cloud is recommended for users needing JSON mode, while Novita and GMI provide viable alternatives for specific integration needs, albeit at a higher cost and lower performance compared to DeepInfra.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Agents | 1 | 4,430 | 1,100 | 236 | -3% |
| Loop engineering | 1 | 53 | 37 | 25 | +18% |
| RAG | 1 | 941 | 216 | 85 | -48% |
| Real-time | 1 | 6,296 | 1,346 | 246 | -2% |
| Reinforcement learning | 1 | 104 | 49 | 23 | -14% |
| Vector Search | 1 | 1,739 | 413 | 146 | -27% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.