NVIDIA open-weights Nemotron 3.5 Lightning โ always-on agents get a fast MoE

Always-on agents burn money on tool calls, not grand plans. Per NVIDIA’s August 11 blog, the company shipped Nemotron 3.5 Lightning: an open ~30B mixture-of-experts model with ~3B active parameters, aimed at the high-volume execution layer inside multi-agent systems. Frontier models still plan. Lightning is built to grind the steps they hand off.
Architecture numbers that matter
Lightning is a hybrid Mamba-2 + MoE + Attention stack with up to 1M tokens of context. MoE routing keeps only a slice of parameters hot per token, so you get denser-model capacity at small-model compute cost. Pre-training used an NVFP4 recipe across more than 20 trillion tokens. Native multi-token prediction plus speculative-decoding drafters (DSpark, DFlash) sit in the release kit. Recommended sampling on the model card: temperature 1.0, top_p 0.95.
Hardware pitch is blunt: serve it on 1ร DGX Spark or 1ร H100, or push it onto RTX PCs, DGX Station, Jetson, and cloud NIMs. Ampere is in via W4A16. Language coverage includes English plus Spanish, French, German, Italian, and Japanese, with coding languages in pre-training. The Nemotron Coalition contributed eval methods, inference software, and datasets. License framing is OpenMDW 1.1 (permissive, commercial use allowed). That is the open-weights story, not another API-only “open” tease.
Claimed speedups on agent tasks
NVIDIA says Lightning delivers up to 4ร faster output versus similar-sized peers and about 30% faster completion on PinchBench-style agent tasks at comparable accuracy. Treat those as vendor benches until third parties publish matching harnesses. The comparison class is other mid-size open MoEs, not Ultra-class planners. The product claim underneath is clearer: long-running agents spend most cycles on code review, tool use, alert triage, and billing answers, not single-shot chat.
Named customizers already in the blog: CrowdStrike (security), Harvey with Trajectory (legal), CodeRabbit with Baseten (code review), plus Lila Sciences and Fastino Labs on domain post-training. That list is marketing, but it maps to the intended job: fine-tune the worker with NeMo on your traces, keep the planner elsewhere.
Where weights ship on Hugging Face
Weights land on Hugging Face under names like NVIDIA-Nemotron-3.5-Lightning-30B-A3B in NVFP4 and BF16, with a base checkpoint for teams that want their own post-train loop rather than a frozen chat endpoint. Also listed: ModelScope, OpenRouter, and build.nvidia.com as a NIM. NVIDIA is publishing an agentic RL dataset (Nemotron-RL-Agentic-Terminal-Pivot) used to sharpen coding-agent behavior. Traceability is the Nemotron habit: as much training data and technique as licensing allows.
How it pairs with Switchyard routing
Same day, NVIDIA open-sourced NeMo Switchyard, a routing library that sends each agent step to the cheapest capable model in a mixed fleet. Internal claims: frontier-level accuracy while cutting task cost to roughly one-third of routing everything to Opus 4.8 alone. Partners named in the same post include LangChain, LiteLLM, Kong, Cognition, and others already wiring routers into gateways and agent shells.
Analysis: Lightning plus Switchyard is NVIDIA arguing the durable agent stack is an ensemble, not one giant chat model. If the PinchBench and cost cuts hold outside NVIDIA’s charts, builders get an open worker they can post-train and a router that stops overpaying for every tool call. If the benches are optimistic theater, you still get usable 30B-A3B weights under OpenMDW and a GitHub router worth trying. Watch independent agent harnesses and how fast the Hugging Face checkpoints show up in production routers.



