NVIDIA NeMo Switchyard Review 2026: Self-Hosted LLM Routing

nvidianemoroutingllmselfhostedagentslitellm

TL;DR: NeMo Switchyard is NVIDIA’s Apache 2.0 routing proxy that decides, per request, whether an agent’s next step goes to a cheap model or an expensive one — using execution signals instead of just prompt classification. It is genuinely open, genuinely clever, and genuinely pre-1.0: the bundled server is labeled “Demo” by NVIDIA’s own stability table. Self-hosters with one Ollama box don’t need it; multi-model stacks with a real cost gradient should watch it closely.

NeMo Switchyard v0.3.0LiteLLM ProxyRouteLLM
Best forStage-aware routing inside agent loopsAggregating 100+ providers behind one endpointResearch-grade cheap/expensive classification
LicenseApache 2.0MIT (proxy core)Apache 2.0
Routing logicExecution signals, LLM classifier, compositeManual rules, fallbacks, load balancingTrained classifier on prompt
MaturityPre-1.0; server component marked “Demo”Production-deployed for yearsResearch project (LMSYS)
The catchNot production-ready by its own docsRouting is static — you write the rulesNo agent/stage awareness

Honest take: run LiteLLM today if you need one endpoint in front of many models. Pin Switchyard v0.3.0 in a homelab experiment if the “route by what the agent is doing” idea maps to a real bill you’re trying to shrink — that idea is the part LiteLLM doesn’t have.

What is NVIDIA NeMo Switchyard?

NeMo Switchyard is an open-source routing layer that sits between an AI agent and a pool of models, and picks which model handles each individual request based on task difficulty, cost, and where the agent is in its workflow. It lives at github.com/NVIDIA-NeMo/Switchyard, shipped as part of NVIDIA’s open-model push alongside the Nemotron line, and reached v0.3.0 as of October 2026.

The pitch is narrow and specific. An agent run is not one request — it’s dozens: planning steps, file edits, retries, recovery from failed tool calls. Sending every one of those to your biggest model wastes money (or VRAM, if you self-host); sending every one to a small model tanks success rates. Switchyard’s answer is per-step routing: steady, mechanical steps go to the efficient model, and exploration or error recovery escalates to the capable one.

Under the hood it’s a Rust core with a Python-facing ecosystem, split into components NVIDIA rates separately for stability: the embeddable library switchyard-libsy is Beta, the HTTP model client is Alpha, and the standalone switchyard-server proxy is explicitly labeled Demo — not for production. Keep that table in mind; it does a lot of work in this review.

Is NeMo Switchyard really open source?

Yes — Apache 2.0, confirmed from the LICENSE file in the repository, copyright NVIDIA Corporation. There is no CUDA lock-in clause, no “NVIDIA hardware only” restriction, and no NIM container requirement. The router itself is a proxy that speaks HTTP to whatever endpoints you configure; it runs on CPU and doesn’t load model weights at all.

That matters because NVIDIA’s open-source record is mixed — plenty of NeMo ecosystem pieces are Apache 2.0, while model weights often carry the NVIDIA Open Model License instead. Switchyard sits on the clean side of that line. If you want the full serving-platform context, our Ollama vs OpenRouter vs Groq vs NVIDIA NIM comparison covers where NIM’s licensing gets murkier.

How does Switchyard decide which model gets the request?

Four routing presets ship in v0.3.0, configured with a single type line in TOML:

  • auto — the default. An execution-stage router with an efficient-first picker and a 0.5 confidence threshold. No extra classifier call, so no added latency or token cost.
  • llm_classifier (the “Task” preset, mode = "capability") — a model judges whether the cheap model can handle the request before routing. Costs one extra small inference per request.
  • stage_router (the “Execution” preset) — routes on recent tool activity and outcome signals: repeated failures or recovery behavior escalate to the capable model, steady edits stay cheap.
  • composite — chains Task and Execution together.

The full catalog also lists plan/execute splitting, escalation, advisor, sub-agent, and random routing, with support varying by gateway. The execution-stage idea is the genuinely novel part. LiteLLM routes on model name, latency, and cost rules you write by hand; RouteLLM classifies the prompt text. Neither looks at how the agent run is going. Switchyard watching tool-call outcomes and escalating when the agent starts flailing is a different — and for coding agents, more honest — signal.

Can Switchyard route to Ollama and other local backends?

Yes, through generic OpenAI-compatible clients — Ollama is not named anywhere in the docs, but it doesn’t need to be. A Switchyard llm_client is defined by a format (openai_chat, openai_responses, or anthropic_messages) plus an arbitrary base_url. Ollama has exposed an OpenAI-compatible API at http://localhost:11434/v1 since early 2024, and vLLM serves the same format, so both slot in as targets.

A realistic homelab config routes mechanical steps to a small local model and hard steps to a bigger one on the same box (or a paid API):

[llm_clients.ollama]
format = "openai_chat"
base_url = "http://localhost:11434/v1"

[targets.weak]
client = "ollama"
model = "qwen3.6-35b-a3b"

[targets.strong]
client = "ollama"
model = "muse-glimmer-30b"

[routes.smart]
id = "switchyard"
type = "auto"
efficient_target = "weak"
capable_target = "strong"

API keys, where a client needs one, are referenced by environment-variable name (api_key_env) so secrets stay out of the TOML. There’s also a forward_auth = true mode that passes each caller’s credential upstream — the docs warn it can leak provider-specific headers, so leave it off on a multi-user box.

One important boundary: Switchyard routes HTTP calls. It does not load or unload models from VRAM. If both your targets live on one 24 GB GPU, Ollama’s own model swapping still governs what’s resident — routing to a cold model eats a load delay Switchyard knows nothing about. Plan VRAM so both targets stay loaded, or point the capable target at a remote endpoint.

How do you install and run switchyard-server?

The server installs from crates.io with a working Rust toolchain — there’s no pip package or official Docker image for it as of v0.3.0, which tells you something about the intended audience:

$ cargo install --locked switchyard-server
$ switchyard-server --config routes.toml --dry-run
$ switchyard-server --config routes.toml --host 127.0.0.1 --port 4000
$ curl http://localhost:4000/health
{"status":"ok"}

The --dry-run flag validates your TOML without opening a socket — use it, because a typo’d target name is otherwise a runtime surprise. Once it’s up, any OpenAI-compatible client talks to it by setting "model": "switchyard" (the route id) against /v1/chat/completions. For embedding instead of proxying, switchyard-libsy is added as a Rust git dependency pinned to the v0.3.0 tag.

Here’s the gotcha that cost me the most time reading this stack: the version pin is not optional. The project is pre-1.0 and says plainly that interfaces move between minors — the NeMo Relay plugin, for instance, requires Relay >=0.8.0, <1.0.0, and the preset catalog already changed shape between 0.1.0 (when early write-ups showed a pip install nemo-switchyard path) and 0.3.0 (where the Rust server is the documented route). Guides from August 2026 already don’t match the current install. If you deploy this, pin the exact tag and treat every upgrade as a migration.

When should you NOT use NeMo Switchyard?

Skip it in any of these situations:

  • You run one model. Routing needs a cost or capability gradient to exploit. A single Ollama instance serving one model has nothing to route between — see when NOT to self-host AI for the broader version of that math.
  • You need production reliability today. NVIDIA’s own stability table calls the server “Demo.” Putting a Demo-grade proxy in front of a team’s inference traffic is how you turn one point of failure into two. LiteLLM’s proxy is years into production hardening (patch it, though — see our LiteLLM CVE-2026-42271 guide).
  • Your “routing” is really provider aggregation. If the goal is one endpoint in front of OpenAI, Anthropic, and a local vLLM — with spend tracking and fallbacks — that’s LiteLLM’s exact job, under MIT.
  • Your agent steps don’t vary in difficulty. Batch summarization, embedding pipelines, template filling: uniform work gains nothing from per-step escalation and just adds a hop.

The cost side is at least cheap to experiment with. The router is a CPU-only Rust binary; a $6/month Vultr VPS runs it comfortably next to the rest of a non-GPU stack. If you want to test a two-tier model pool without buying a second GPU, a rented RTX 3090 from Vast.ai (from ~$0.07/hr, floating market pricing) as the “capable” target against a local small model is the cheapest way to find out whether stage routing actually cuts your token spend. The hardware side of running two models resident at home is runaihome.com’s territory.

FAQ

Does NeMo Switchyard require NVIDIA GPUs or NIM containers? No. The router is hardware-agnostic — it’s a CPU-bound proxy speaking OpenAI Chat, OpenAI Responses, and Anthropic Messages formats to any base_url you configure. NIM endpoints work as targets, but so does a Raspberry Pi running llama.cpp’s server, and nothing in the Apache 2.0 license restricts non-NVIDIA use.

Can Claude Code or other coding agents use it without modification? Yes, in proxy mode — that’s the design center. Because Switchyard translates between Anthropic Messages and OpenAI Chat formats, an agent that speaks one API can be routed to backends that speak the other. Point the agent’s base URL at the Switchyard server and set the model to your route id.

Is Switchyard better than LiteLLM? Different jobs. LiteLLM is a mature aggregation proxy with manual routing rules; Switchyard is an early-stage router with automatic, execution-aware model selection. They’re not even mutually exclusive — LiteLLM is listed among Switchyard’s gateway integrations, alongside OpenRouter (where it appears as the nvidia/switchyard model) and NeMo Relay.

Sources

Was this article helpful?

What self-hosting actually costs

Real cost breakdowns for self-hosted AI: hardware floors, power, maintenance hours, and the honest comparison against paying for it. No spam, unsubscribe anytime.