Home / Companies / Kong / Blog / Post Details
Content Deep Dive

Intelligent Model Routing: NVIDIA NeMo Switchyard & Kong AI Gateway

Blog post from Kong

Post Details
Company
Date Published
Author
Harish Madhavan and Grajesh Chandra
Word Count
1,383
Company Posts That Month
4
Language
English
Hacker News Points
-
Post removed?
No
Summary

Intelligent LLM routing can reduce costs and improve latency by selecting models according to task complexity and quality requirements, but placing routing logic directly in production request paths creates security, reliability, and governance challenges. Kong proposes separating model selection from traffic management by pairing NVIDIA’s open-source NeMo Switchyard, which provides configurable model-selection algorithms, with Kong AI Gateway, which retains responsibility for routing, credentials, guardrails, PII masking, rate limits, auditing, caching, failover, and multi-environment deployment. In this architecture, Kong sends request information to Switchyard for a model-target decision, applies organizational policies, and routes the request, while configurable fallbacks preserve availability if the decision service is unavailable and session persistence can reduce repeated evaluations in multi-turn conversations. The approach is intended to let ML teams tune selection policies independently while platform teams maintain a centralized security and governance boundary for model calls as well as related APIs, MCP servers, and event streams. In a preliminary OpenThoughts-TBLite benchmark, raising a Switchyard routing-confidence threshold reduced escalation to a frontier model from 85% to 17% and lowered cost per completed task by 43.7% without reducing completion rates, though further evaluation is planned.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.