Home / Companies / Deepinfra / Blog / Post Details
Content Deep Dive

Qwen3.8-27B API Provider Benchmarks: Speed & Cost

Blog post from Deepinfra

Post Details
Company
Date Published
Author
Deep
Word Count
2,641
Company Posts That Month
11
Language
English
Hacker News Points
-
Post removed?
No
Summary

DeepInfra’s review compares API providers for Alibaba’s Qwen3.8-27B, a 27.78-billion-parameter open-weight reasoning and vision-language model released under Apache 2.0 in August 2026. The model supports text, image, and video inputs, produces text output, offers a 262K-token native context window extendable to 1 million tokens, and is positioned for coding, agentic workflows, data analysis, and visual reasoning. Citing Artificial Analysis benchmarks, the review presents DeepInfra as a cost-focused option with promotional pricing of $0.15 per million input tokens and $1.875 per million output tokens, though its approximately 29-token-per-second throughput is slower than competitors. Multiverse Computing and Crusoe are identified as the fastest options at roughly 191–192 output tokens per second, while Multiverse also has the lowest reported time to first token at 0.66 seconds; OpenRouter emphasizes multi-provider routing and extended context support, and Alibaba Cloud offers first-party access with slower performance and higher pricing. The review attributes Qwen3.8-27B’s strong benchmark results to its hybrid-attention dense architecture and adjustable reasoning effort, while noting that verbose reasoning can increase output costs and latency.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 6 No monthly metrics for this publish month.
AI Agents 1 No monthly metrics for this publish month.
Real-time 1 No monthly metrics for this publish month.
Vector Search 1 No monthly metrics for this publish month.
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.