Model details
Switchyard
NVIDIA NeMo Switchyard is an open-source routing library designed to sit between inference endpoints and the models they serve, intelligently directing each prompt across a heterogeneous mix of open, proprietary, and NVIDIA models. Rather than sending all traffic to a single model, the library can be configured to route requests based on cost, latency, or output quality, giving enterprises a way to blend cheaper local models with more expensive frontier systems within the same application. It was announced alongside Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model in the Nemotron 3 family aimed at long-running agentic AI workloads.
The practical motivation behind NeMo Switchyard is managing the soaring infrastructure costs and unpredictable returns that can stall enterprise AI rollouts: by letting developers route only the prompts that need a stronger model to a premium system while keeping simpler queries local or on smaller open models, the library aims to make AI spend more manageable without rewriting existing applications. NVIDIA positions it as part of a broader push toward customizable open systems that can run across PCs, workstations, data centers, and the cloud, complementing rather than replacing the underlying models themselves.
Quick Info
Powered by- Provider
- OpenRouter
- Model key
- nvidia/switchyard
- Input modalities
- Output modalities
- Capabilities
- Base catalog fields only
Limits
- Context window
- 1,000,000 tokens
Latest news about Switchyard
No articles yet. Fetch the latest news to show it here.
