Who Decides Which Model Runs? NVIDIA Would Like a Say
NVIDIA announced Nemotron 3.5 Lightning, a 30B open mixture-of-experts model with about 3B active parameters, and NeMo Switchyard, an open source routing library for agent workflows. NVIDIA says Lightning delivers 4x throughput versus comparable models and up to 30% faster agentic benchmark completion. NVIDIA also claims routed task costs drop to about one-third with completion rates broadly unchanged.
How this was made

The 30-second read
Why it matters
If developers adopt NeMo Switchyard with Lightning, NVIDIA could benefit from increased demand for its inference stack (NIM, DGX Spark/Station, Jetson/RTX) and from software ecosystem lock-in via gateways.
Market read
Traders may view this as a software and efficiency play that supports NVIDIA’s AI platform narrative, but it is not accompanied by financial guidance or confirmed large-scale deployments.
What to watch
Routing introduces new observability, attribution, and compliance overheads; without strong tooling and adoption, the economic advantage may erode in regulated or production settings.
Background
The article frames NVIDIA’s strategy as extending open model efforts into the routing layer that decides which model runs at each step of an agentic workflow.
Ticker impact
NVIDIA announced Nemotron 3.5 Lightning and NeMo Switchyard, claiming 4x throughput and up to 30% faster agentic benchmarks.
Near-term sentiment likely positive for NVDA AI platform positioning, but magnitude depends on whether customers validate the cost and routing claims.
This is a product and ecosystem announcement with quantified performance/cost claims, but the article provides no new financial guidance or confirmed customer adoption beyond early-access partners.
Market effects
Could increase competitive focus on inference efficiency, model routing, and on-device/edge deployment for agentic AI stacks.
No clear regional macro linkage; impact is primarily global AI infrastructure and developer tooling.
If validated, routing and local execution economics may influence how enterprises procure AI compute worldwide.
Counterpoint
Benchmark and cost claims may not generalize to broader, non-research customer data and operational constraints, limiting commercial impact.
Key entities
- productNVIDIA Nemotron 3.5 Lightning
A 30B open mixture-of-experts model with ~3B active parameters, positioned for always-on agent call throughput.
- softwareNeMo Switchyard
An open source model routing library that selects models per agent workflow step and integrates with OpenRouter, LiteLLM, and Kong.
- customer_partnerCrowdStrike
Early post-training partner cited for improved benign recall after customization.
- customer_partnerCodeRabbit
Early post-training partner cited for coding router improvements and a stated build cost/time.
- customer_partnerHarvey with Trajectory
Early post-training partner cited for legal task completion gains.



