Step 3.5 Flash API Benchmarks: Latency, Throughput & Cost
Blog post from Deepinfra
Step 3.5 Flash is an advanced open-weights reasoning model launched by StepFun in February 2026, employing a Mixture of Experts architecture with 196 billion parameters, and is distinguished by having only 11 billion active parameters per token during inference. This model achieves high performance with a score of 38 on the Artificial Analysis Intelligence Index and supports a large context window of 256k tokens, facilitating extensive reasoning capabilities and structured outputs via JSON mode. DeepInfra is highlighted as the optimal provider for deploying Step 3.5 Flash due to its industry-leading latency of approximately 0.32 seconds, competitive pricing of $0.10 per million input tokens and $0.30 per million output tokens, and support for full JSON Mode and Function Calling. The model's verbose nature, producing 200 million tokens during evaluations, makes cost efficiency pivotal, and DeepInfra's infrastructure offers the best balance for real-time applications. SiliconFlow is recommended for high-throughput batch tasks, while StepFun provides a reliable baseline for non-interactive applications, and OpenRouter ensures API redundancy for enterprise needs.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Real-time | 4 | 6,296 | 1,346 | 246 | -2% |
| LLM | 2 | 5,932 | 1,046 | 223 | -2% |
| AI Agents | 1 | 4,430 | 1,100 | 236 | -3% |
| Vector Search | 1 | 1,739 | 413 | 146 | -27% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.