Home / Companies / Deepinfra / Blog / Post Details
Content Deep Dive

Qwen3.5 397B A17B API Benchmarks: Latency, Throughput & Cost

Blog post from Deepinfra

Post Details
Company
Date Published
Author
Deep
Word Count
2,094
Company Posts That Month
34
Language
English
Hacker News Points
-
Post removed?
No
Summary

Alibaba Cloud's Qwen3.5 397B A17B, released in February 2026, is a multimodal foundation model that integrates text and vision capabilities in a unified architecture. Featuring a hybrid Mixture-of-Experts design with 397 billion total parameters, this model offers significant improvements in latency and cost efficiency through its sparse activation approach. The model supports a wide array of functionalities, including reasoning and non-reasoning modes, a 262k token context window, and compatibility with 201 languages and dialects. Among various providers, DeepInfra (FP8) is recommended for production deployment due to its lowest latency (0.67 seconds), competitive pricing ($1.25 per 1M tokens), and high throughput (137.9 tokens per second). Clarifai stands out for maximum throughput, while Eigen AI is favored for structured data extraction. Although Alibaba Cloud provides full feature support and first-party hosting, its latency and throughput are outpaced by third-party alternatives. The analysis indicates that DeepInfra offers the best overall value for deploying the Qwen3.5 397B A17B model at scale, while Clarifai and Eigen AI serve specialized needs like high-speed generation and structured data handling, respectively.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 5 5,932 1,046 223 -2%
AI Agents 1 4,430 1,100 236 -3%
Real-time 1 6,296 1,346 246 -2%
Vector Search 1 1,739 413 146 -27%
Voice AI 1 2,379 221 38 -3%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.