As artificial intelligence shifts from single-turn chatbots to autonomous agents that run continuously, NVIDIA is expanding its open model lineup to meet the moment. On August 11, 2026, the company unveiled Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model it calls the highest-efficiency model in its class for long-running agentic AI workloads. The release, detailed on the NVIDIA blog by Kari Briski, arrives alongside NeMo Switchyard, an open source routing library designed to direct each task to the most capable and cost-effective model automatically. For the latest AI industry coverage, follow along as the agentic AI stack takes shape.
A Model Built for Always-On Agents
Modern agentic systems increasingly operate as what NVIDIA describes as "systems of models" or model ensembles. A frontier reasoning model such as Nemotron 3 Ultra or GPT-5.6 might plan and orchestrate a workflow, while smaller specialized models handle targeted tasks like code review, tool use, security alert monitoring, or answering billing questions. Nemotron 3.5 Lightning is purpose-built for that second tier: high-volume, specialized tasks that power always-on agents.
According to NVIDIA, the model delivers up to 4x faster output speed, leading to roughly 30% faster agentic task completion compared with other models in its class. NVIDIA's own PinchBench benchmarks reportedly demonstrate that Nemotron 3.5 Lightning maintains frontier-level accuracy while completing tasks more quickly than its peers. The model was developed with contributions from the Nemotron Coalition, whose members provided evaluation methodologies, inference software, and datasets to help advance the model.
Because it is open and customizable, organizations can post-train Nemotron 3.5 Lightning with NVIDIA NeMo on their own domain data, tools, and workflows to improve accuracy for specialized tasks. NVIDIA is also releasing Nemotron-RL-Agentic-Terminal-Pivot, an agentic reinforcement learning dataset used to post-train the model for coding agent capabilities.
Running Anywhere, From Edge to Cloud
One of the central selling points of Nemotron 3.5 Lightning is deployment flexibility. The model can run on local AI systems — including NVIDIA RTX PCs, NVIDIA DGX Spark, NVIDIA DGX Station, and NVIDIA Jetson — letting organizations maximize existing infrastructure investments. It can also scale across edge AI devices, NVIDIA RTX PRO workstations, data centers, and cloud environments for enterprise use cases. That flexibility matters for organizations handling sensitive data that must remain on-premises.
As with every Nemotron launch, NVIDIA says it publishes as much of the training data and techniques as licensing permits, which allows for traceability, auditing, and training of other models. This commitment to transparency has become a competitive differentiator as enterprises weigh open models against proprietary alternatives that offer less visibility into how they were built.
NeMo Switchyard: Smart Routing for Systems of Models
If Nemotron 3.5 Lightning is the workhorse, NeMo Switchyard is the dispatcher. The open source library routes prompts to the most capable and efficient model for each step of an agent workflow automatically, based on specific needs. Some models are better for coding, some for reasoning, some for lightweight tasks, and some are optimized to run locally for greater privacy and efficiency. Relying on a single default model can mean overspending or sacrificing quality; managing routing manually, meanwhile, becomes integration work that can slow down a deployment.
Agent application developers can tune or modify the router with different routing algorithms to match their priorities, such as quality, latency, and cost requirements. NVIDIA's internal benchmarks show that NeMo Switchyard maintains frontier-level accuracy while reducing task completion cost to nearly one-third of Opus 4.8 alone — a striking efficiency gain for organizations running large numbers of agent tasks.
The company highlighted a string of partner integrations. Boomi reported 100% domain-routing accuracy and sent 59% of traffic to a 5x faster fine-tuned model. Cadence improved efficiency by 9.9% on a formal verification use case. Cognition integrated Switchyard into Devin Desktop, achieving near-frontier performance on FrontierCode Main while cutting mean cost by 28%. LangChain reported 74% lower cost across 145 multi-turn Deep Agents tasks by routing only 7% of calls to a frontier model. LiteLLM is adding NeMo Switchyard as a plug-in to its proxy layer, and Kong now delivers routing natively through the Kong AI Gateway. Notably, Nous Research integrated NeMo Switchyard into Hermes to give developers an easy-to-configure routing system for agent efficiency.
Industry Adoption and the Bigger Picture
AI leaders across industries are already customizing Nemotron 3.5 Lightning for their workloads. CrowdStrike is applying it to cybersecurity, Harvey is using it for legal services, and CodeRabbit is deploying it for code review. Lila Sciences is helping improve reasoning capabilities for agentic tasks across physical and life sciences, while Fastino Labs customized the model and reported leading accuracies for software development, finance, and healthcare workloads.
The release underscores a broader industry trend: as agentic AI matures, no single model can do everything well at an acceptable cost. The future belongs to ensembles — frontier planners coordinating fleets of specialized workers — and the tooling that ties them together. NVIDIA's bet is that open models and intelligent routing will let enterprises assemble those systems without locking themselves into a single provider's stack.
Stay Ahead of AI
NVIDIA's Nemotron 3.5 Lightning and NeMo Switchyard signal where agentic AI is heading: faster, cheaper, more transparent, and deployable anywhere. To keep up with the model releases and enterprise tools reshaping the landscape, explore our breaking AI news.
Read more AI news →