Home / Companies / Deepinfra / Blog / Post Details
Content Deep Dive

Qwen3.5 35B A3B API Benchmarks: Latency, Throughput & Cost

Blog post from Deepinfra

Post Details
Company
Date Published
Author
Deep
Word Count
1,201
Company Posts That Month
34
Language
English
Hacker News Points
-
Post removed?
No
Summary

Qwen3.5 35B A3B, a vision-language model launched by Alibaba Cloud in 2026, integrates Gated Delta Networks with a sparse Mixture-of-Experts model to achieve higher inference efficiency with 35 billion parameters, activating only 3 billion per token. Offering a 262k token context window, tool calling, dual thinking modes, and support for 201 languages, the model is available through various providers under the Apache 2.0 license. DeepInfra (FP8) emerges as the preferred API for interactive applications due to its unmatched initial response time of 0.60 seconds and fastest end-to-end performance of 14.86 seconds, while GMI (FP8) excels in throughput with 190 tokens per second. Novita provides a balanced option with competitive pricing and performance, and Alibaba Cloud offers first-party support and extended features via the Qwen3.5-Flash hosted API. The choice of provider depends on specific application needs, balancing between speed, cost, and throughput.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 4 6,296 1,346 246 -2%
LLM 3 5,932 1,046 223 -2%
AI Agents 1 4,430 1,100 236 -3%
Vector Search 1 1,739 413 146 -27%
Voice AI 1 2,379 221 38 -3%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.