Home / Companies / Weave / Blog / Post Details
Content Deep Dive

The top 10 models our router actually uses, and what each one is for

Blog post from Weave

Post Details
Company
Date Published
Author
-
Word Count
1,247
Company Posts That Month
6
Language
English
Hacker News Points
-
Post removed?
No
Summary

Over a 30-day period, Weave routed about 235,000 requests across roughly 40 AI models using complexity clusters—fast, balanced, high, and maximum—alongside hard-pinned utility tasks, user-selected models, and legacy routing rules. DeepSeek-V4-Flash handled the largest share of requests by serving inexpensive fast tasks, subagent bootstraps, and tool-result follow-ups, while Gemini 3.1 Flash Lite was used exclusively for pinned classification, title generation, and health checks. Minimax-M3 dominated routine balanced tasks under a newer XGBoost-based selector, whereas Claude Sonnet 5 led high-complexity coding work and consumed the most tokens. The maximum tier, including Claude Opus 5, Claude Fable 5, GPT-5.5, and GPT-5.6-Sol, handled fewer but more costly and often longer-running sessions, with premium models retaining users across many turns. The data shows a large cost imbalance: the cheapest models processed about 38% of requests for roughly 1% of spending, while Claude models represented about 89% of total spend despite handling about 38% of requests, partly offset by higher cache-read rates. Model usage changes quickly as models are introduced, replaced, or gain learned routing preference, illustrating that the router continuously adapts rather than relying on fixed assignments.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.