Intelligent Model Routing: NVIDIA NeMo Switchyard & Kong AI Gateway
Blog post from Kong
Intelligent LLM routing can reduce costs and improve latency by selecting models according to task complexity and quality requirements, but placing routing logic directly in production request paths creates security, reliability, and governance challenges. Kong proposes separating model selection from traffic management by pairing NVIDIA’s open-source NeMo Switchyard, which provides configurable model-selection algorithms, with Kong AI Gateway, which retains responsibility for routing, credentials, guardrails, PII masking, rate limits, auditing, caching, failover, and multi-environment deployment. In this architecture, Kong sends request information to Switchyard for a model-target decision, applies organizational policies, and routes the request, while configurable fallbacks preserve availability if the decision service is unavailable and session persistence can reduce repeated evaluations in multi-turn conversations. The approach is intended to let ML teams tune selection policies independently while platform teams maintain a centralized security and governance boundary for model calls as well as related APIs, MCP servers, and event streams. In a preliminary OpenThoughts-TBLite benchmark, raising a Switchyard routing-confidence threshold reduced escalation to a frontier model from 85% to 17% and lowered cost per completed task by 43.7% without reducing completion rates, though further evaluation is planned.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.