← Back to Blog

NVIDIA Nemotron 3.5 Lightning: The 670 tok/s Open Execution Layer for Always-On Agents

Category: Model Releases · 2026-08-17 · ~5 min read · Verified metrics & benchmark data

The Verdict First

NVIDIA has released Nemotron 3.5 Lightning (August 11, 2026), the first model in the Nemotron 3.5 family and the successor to Nemotron 3 Nano 30B A3B. It is built for the "orchestrator-executor" pattern: a frontier brain like Nemotron 3 Ultra plans, and Lightning executes the high-volume, high-frequency steps. With 31.6B total parameters but only 3.6B active per forward pass, a 1M token context window, and throughput of nearly 670 tokens per second on a DeepInfra NVFP4 endpoint, it is one of the fastest open-weights reasoning models available right now — under the permissive OpenMDW-1.1 license.

31.6B
Total Params (3.6B Active)
1M
Token Context Window
24
AA Intelligence Index
~670 tok/s
NVFP4 Throughput (DeepInfra)

Architecture: A Hybrid Mamba-Transformer MoE

Nemotron 3.5 Lightning is a Mixture-of-Experts (MoE) model with a hybrid Mamba-Transformer architecture. The total parameter count is 31.6B, but only 3.6B parameters are active during any single forward pass — the design goal being high intelligence density without the compute cost of a dense model of the same size. The model is text-only, optimized for reasoning, and carries a 1 million token context window.

Benchmarks: The +9 Point Jump

On the Artificial Analysis Intelligence Index, Nemotron 3.5 Lightning scored 24 — a +9 point jump over its predecessor Nemotron 3 Nano at 15. That matches OpenAI's gpt-oss-120b while using only about a quarter of the total parameters, and it trails only Nemotron 3 Super (26), which is roughly 4x its size. On Terminal-Bench v2.1 the model jumped from 7% to 24%, and on GDPval-AA v2 it gained +334 ELO, outperforming gpt-oss-120b.

Model AA Intelligence Index Relative Size Primary Role
Nemotron 3.5 Lightning 24 31.6B total / 3.6B active High-Volume Execution
Nemotron 3 Super 26 ~4x Lightning's size Frontier Orchestration
gpt-oss-120b (OpenAI) 24 ~4x Lightning's total params Matched by Lightning
Nemotron 3 Nano 30B A3B 15 Predecessor Replaced by Lightning

💡 Key Finding: The biggest agentic leap is Terminal-Bench v2.1: 7% → 24% in a single generation, plus a +334 ELO gain on GDPval-AA v2 — NVIDIA is optimizing specifically for tool-use loops, not chat.

NVFP4: Near-Lossless 4-Bit Shipping

Lightning ships with near-lossless NVFP4 quantization as the standard distribution format. The NVFP4 weights themselves measured 24 on the Intelligence Index — identical to the model's headline score — which means the quantized artifact you actually deploy gives up nothing on the benchmark. This is what makes single-GPU and edge deployments practical: extreme efficiency with no measured intelligence loss. The NVFP4 build is available on Hugging Face, with high-speed endpoints such as DeepInfra serving nearly 670 tokens per second.

The Ecosystem: Switchyard Routing and NemoClaw

Lightning is designed to pair with the new NVIDIA NeMo Switchyard library, which routes tasks to the most cost-effective model in a stack — expensive frontier calls for planning, cheap Lightning calls for execution. For production agents, it is supported by the NVIDIA NemoClaw open-source stack for security and management of always-on operations, and the release notes list compatibility with OpenClaw and Hermes Agent frameworks. The practical effect: token costs drop because repetitive reasoning work is offloaded from frontier models.

Where Lightning Shines

Getting Started

⚡ Explore More AI Model Deep Dives

Read comprehensive benchmark reviews and analysis on CLAW Blog

Fact Verification & Sources

This technical summary was compiled exclusively using verified public data points from the following release documentation:

⚡ OpenAdapter Readers get 20% off — invite code BDPBCR3R ◉ Z.ai Coding Plan Readers get 10% off — invite code R0K78RJKNW
R
Analyzed for CLAW

Live analysis published on claw.rommark.dev on Aug 17, 2026. Data grounded exclusively in official release benchmarks and verified technical specifications.