How many of your agent's calls actually need a frontier model?
Blog post from LangChain
NVIDIA’s open-source NeMo Switchyard routes agent workflow calls between cheaper and frontier language models to reduce spending while reserving more capable models for difficult tasks. In tests on 145 multi-step Deep Agents tasks, an escalation configuration using Nemotron 3.5 Lightning as the default model, Claude Opus 4.8 for escalated sessions, and a small judge model sent only 7% of calls to Opus, cutting cost by 74% relative to Opus alone while retaining 93% of its accuracy, though it remained six percentage points less accurate. The results showed that the cheaper model handled 93% of calls, while the judge accounted for 21.2% of routed spending and frontier-model escalation rates created substantial run-to-run cost variability. The authors emphasize that routing is most useful when teams need high-end capability for unpredictable hard requests, rather than when minimum cost is the sole priority, and propose a formula comparing judge cost with the price difference between models to determine whether routing can save money. They also caution that routing adds latency, works best for multi-turn workloads, and should be evaluated against an organization’s own traffic because the benchmark was relatively saturated and may not generalize.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 8 | 5,068 | 1,020 | 229 | -34% |
| Observability | 1 | 3,175 | 737 | 186 | -24% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.