NVIDIA Releases Nemotron 3.5 Lightning for Local AI Agents
NVIDIA announced on August 11, 2026 that it has expanded its Nemotron 3 model family with Nemotron 3.5 Lightning, an open 30B mixture-of-experts model built for always-on agents. Everything below comes from NVIDIA's own blog post, which is a running August series on local AI, with new entries added over the coming weeks.
Transcript
NVIDIA released Nemotron three point five Lightning today, an open thirty billion parameter model built for always-on local agents.
It is a customizable mixture-of-experts model for always-on agents. NVIDIA says it delivers up to four times faster token generation than open models in its class.
Because the weights are open, developers can fine-tune it for their own work. It runs locally through vLLM, Ollama, llama.cpp and LM Studio, from RTX PCs to data centers.
NVIDIA released NeMo Switchyard, an open source library routing each agent step to the best model. NVIDIA's internal benchmarks show one third the cost of Opus four point eight.
Inboxsmith helps small businesses handle calls and messages so nothing gets missed. Please like and subscribe for more news.
Sources
Every claim in this video comes from the top ranking coverage of this topic. The claims and where each one came from:
- Today, NVIDIA expanded its Nemotron 3 model family with Nemotron 3.5 Lightning, a customizable open 30B mixture-of-experts (MoE) model for always-on agents.(NVIDIA's official announcement)
- Nemotron 3.5 Lightning delivers up to 4x faster token generation and 30% faster time to completion compared to open models in its class.(NVIDIA's official announcement)
- And because Nemotron 3.5 Lightning is open weights, AI enthusiasts and developers can fine-tune it with their own examples to better match specific tasks, interests and workflows.(NVIDIA's official announcement)
- NVIDIA collaborated with vLLM, Ollama, llama.cpp and LM Studio to provide the best local deployment experience for Nemotron 3.5 Lightning models.(NVIDIA's official announcement)
- Nemotron 3.5 Lightning runs locally on NVIDIA RTX PCs, NVIDIA DGX Spark and OEM GB10 systems, and NVIDIA Jetson, and scales up to RTX PRO workstations, NVIDIA DGX Station and GB300 deskside systems, data centers and cloud environments.(NVIDIA's official announcement)
- NVIDIA NeMo Switchyard, an open source routing library, automatically directs each step of an agent workflow to the best-fit model based on accuracy, speed and cost.(NVIDIA's official announcement)
- Internal benchmarks show that NeMo Switchyard, by routing each step across a system of models, helped maintain frontier-level task completion while reducing benchmark completion cost to roughly one-third of Opus 4.8 alone.(NVIDIA's official announcement)
