Home / Companies / Deepinfra / Blog / Post Details
Content Deep Dive

Qwen3.5 0.8B API Benchmarks: Latency, Throughput & Cost

Blog post from Deepinfra

Post Details
Company
Date Published
Author
han
Word Count
1,312
Company Posts That Month
34
Language
English
Hacker News Points
-
Post removed?
No
Summary

Qwen3.5 0.8B, a model in Alibaba Cloud's Qwen3.5 Small Model Series, focuses on delivering high-quality performance on edge devices and mobile phones while maintaining low memory and battery use. This model, designed with an Efficient Hybrid Architecture that includes Gated Delta Networks and sparse Mixture-of-Experts, supports a context window of 262,000 tokens and provides native multimodal capabilities through early fusion training. Released under the Apache 2.0 license for commercial use, it supports 201 languages and dialects, extended reasoning, and function calling for agentic workflows. DeepInfra is the sole benchmarked provider for Qwen3.5 0.8B, offering superior performance metrics like a median TTFT of 0.37 seconds, a throughput of 403.5 tokens per second, and a cost-effective blended price of $0.02 per million tokens. These features, combined with its robust support for JSON mode and function calling, make it an ideal choice for both real-time and batch processing applications.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 3 5,932 1,046 223 -2%
Real-time 3 6,296 1,346 246 -2%
AI Agents 2 4,430 1,100 236 -3%
RAG 2 941 216 85 -48%
AI Model Fine-tuning 1 420 130 55 -54%
Vector Search 1 1,739 413 146 -27%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.