NVIDIA has released Nemotron 3.5 Lightning (August 11, 2026), the first model in the Nemotron 3.5 family and the successor to Nemotron 3 Nano 30B A3B. It is built for the "orchestrator-executor" pattern: a frontier brain like Nemotron 3 Ultra plans, and Lightning executes the high-volume, high-frequency steps. With 31.6B total parameters but only 3.6B active per forward pass, a 1M token context window, and throughput of nearly 670 tokens per second on a DeepInfra NVFP4 endpoint, it is one of the fastest open-weights reasoning models available right now — under the permissive OpenMDW-1.1 license.
Nemotron 3.5 Lightning is a Mixture-of-Experts (MoE) model with a hybrid Mamba-Transformer architecture. The total parameter count is 31.6B, but only 3.6B parameters are active during any single forward pass — the design goal being high intelligence density without the compute cost of a dense model of the same size. The model is text-only, optimized for reasoning, and carries a 1 million token context window.
On the Artificial Analysis Intelligence Index, Nemotron 3.5 Lightning scored 24 — a +9 point jump over its predecessor Nemotron 3 Nano at 15. That matches OpenAI's gpt-oss-120b while using only about a quarter of the total parameters, and it trails only Nemotron 3 Super (26), which is roughly 4x its size. On Terminal-Bench v2.1 the model jumped from 7% to 24%, and on GDPval-AA v2 it gained +334 ELO, outperforming gpt-oss-120b.
| Model | AA Intelligence Index | Relative Size | Primary Role |
|---|---|---|---|
| Nemotron 3.5 Lightning | 24 | 31.6B total / 3.6B active | High-Volume Execution |
| Nemotron 3 Super | 26 | ~4x Lightning's size | Frontier Orchestration |
| gpt-oss-120b (OpenAI) | 24 | ~4x Lightning's total params | Matched by Lightning |
| Nemotron 3 Nano 30B A3B | 15 | Predecessor | Replaced by Lightning |
💡 Key Finding: The biggest agentic leap is Terminal-Bench v2.1: 7% → 24% in a single generation, plus a +334 ELO gain on GDPval-AA v2 — NVIDIA is optimizing specifically for tool-use loops, not chat.
Lightning ships with near-lossless NVFP4 quantization as the standard distribution format. The NVFP4 weights themselves measured 24 on the Intelligence Index — identical to the model's headline score — which means the quantized artifact you actually deploy gives up nothing on the benchmark. This is what makes single-GPU and edge deployments practical: extreme efficiency with no measured intelligence loss. The NVFP4 build is available on Hugging Face, with high-speed endpoints such as DeepInfra serving nearly 670 tokens per second.
Lightning is designed to pair with the new NVIDIA NeMo Switchyard library, which routes tasks to the most cost-effective model in a stack — expensive frontier calls for planning, cheap Lightning calls for execution. For production agents, it is supported by the NVIDIA NemoClaw open-source stack for security and management of always-on operations, and the release notes list compatibility with OpenClaw and Hermes Agent frameworks. The practical effect: token costs drop because repetitive reasoning work is offloaded from frontier models.
This technical summary was compiled exclusively using verified public data points from the following release documentation: