NVIDIA releases NeMo Switchyard to route agents across AI models

Always-on agents are becoming systems of models, and pinning every step to one frontier SKU is how you burn the budget. Per NVIDIA’s Nemotron Lightning / Switchyard blog, the company open-sourced NeMo Switchyard: a routing library that sends each agent step to the most capable, suitable model across a mix of open, proprietary, and NVIDIA models โ without rewriting the app.
That is the product cut. Lightning is the specialized open weights. Switchyard is the traffic cop that decides when you actually need them.
What Switchyard routes per step
Modern agent stacks already split labor: a frontier reasoner plans, while smaller models handle code review, tool calls, alert triage, or billing FAQs. Switchyard automates that split at request time. Developers build a router tuned to their priorities, then each prompt or step lands on the model that fits โ quality, latency, cost, or local privacy โ instead of a single default that is either too expensive or too weak.
NVIDIA’s framing is blunt. One default model means overspend or lost quality. Manual routing means integration work that slows deployment. Switchyard turns routing into a library concern so the agent graph stays stable while the model mix changes underneath.
Integrations with LangChain and LiteLLM
The launch is as much about ecosystem glue as about the router itself. LangChain reports 74% lower cost across 145 multi-turn Deep Agents tasks by routing only 7% of calls to a frontier model, at a 6% accuracy tradeoff. LiteLLM is adding Switchyard as a plug-in in its proxy layer so teams keep their existing stack. Kong ships routing natively through Kong AI Gateway. Nous Research integrated it into Hermes. Cognition wired a staged Switchyard router into Devin Desktop for NVIDIA internal use, claiming near-frontier FrontierCode Main performance with 28% lower mean cost versus sending everything to one frontier model.
Other named partners include Boomi, Cadence, Classmethod, Ramp, and Siemens. Treat the board as vendor-reported case notes, not a third-party audit.
Cost and latency knobs
Switchyard’s selling point is tunable tokenomics. Developers can modify routing algorithms to favor quality, latency, or cost. NVIDIA’s internal benchmarks claim frontier-level accuracy while cutting task completion cost to nearly one-third of Opus 4.8 alone. Partner snippets echo the theme: Boomi sending 59% of traffic to a 5ร faster fine-tuned model with 21% lower later-turn latency; Classmethod seeing 27% cost reduction at maintained quality; Ramp matching frontier performance while cutting costs 58% and runtime 33% on a SWE-Bench slice.
Big savings can buy a small accuracy haircut โ the LangChain caveat. Teams that need every step at frontier grade will still pin. Everyone else gets a dial instead of a religion.
How it complements Nemotron Lightning
Switchyard ships beside Nemotron 3.5 Lightning, a 30-billion-parameter MoE open model aimed at high-volume specialized agent tasks โ up to 4ร faster output and about 30% faster agentic completion versus peers in its class, per NVIDIA, with post-training via NeMo on private tools and data. Lightning is the cheap, customizable specialist inside a multi-model system. Switchyard is how that specialist gets traffic without hard-coding routes.
Together they argue for ensembles you control: frontier planner here, Lightning-class worker there, proprietary model where contracts demand it, local RTX/DGX where privacy wins. Switchyard is on GitHub; Lightning lands on Hugging Face, ModelScope, OpenRouter, and build.nvidia.com as NIM.
Analysis: NVIDIA is productizing the same insight IDE routers already sell โ the model is not the agent; the router is. If Switchyard’s partner savings survive outside launch blogs, defaulting every agent turn to Opus-class pricing starts looking like negligence. Watch whether the GitHub library stays framework-portable, or quietly steers traffic toward Nemotron SKUs whenever “best” is ambiguous.




[…] previously covered NVIDIA NeMo Switchyard, a library for routing work between models. Singtel packages model access and operational controls […]