Model routing has to account for the warm cache
Blog post from Factory
Model routing for AI coding agents should account for full session economics rather than token prices alone, particularly the value of warm cached context, prior tool results, and accumulated decisions that may be costly to reprocess after switching models. Factory reports separate production findings of 58% lower aggregate cost and 49-second median routed-session latency versus 81 seconds for frontier-pinned sessions, while its current documentation cites 43% savings against top-tier pricing; both figures use distinct baselines and are not guarantees. Routing decisions can use task signals such as test outcomes, recurring errors, and context needs to determine when inexpensive models are appropriate and when stronger models should remain active, as illustrated by a 67-turn Prisma migration that could not safely step down. Focused secondary workers may preserve a parent session’s warm context but add their own costs and integration overhead. Enterprises evaluating routing should measure billed costs, cache activity, chosen models, failed attempts, reruns, accepted changes, elapsed time, and cost per successful task under consistent acceptance criteria, while also confirming that model availability, deployment requirements, and enterprise controls fit their environment.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Coding Assistant | 1 | 341 | 115 | 55 | -77% |
| Cost per task | 1 | 10 | 5 | 5 | -84% |
| Local AI | 1 | 15 | 4 | 3 | -94% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.