← Back to Blog
⚡ 10% OFF Get 10% off any Z.AI Coding Plan — covers GLM 5.1, GLM 5 Turbo, and GLM 4.7 → Claim discount
May 03, 2026 8 min read

NIST Evaluation: DeepSeek V4 Pro Trails US Models by 8 Months

The most detailed US government evaluation of Chinese AI finds DeepSeek V4 Pro is the most capable Chinese model but significantly behind US frontier systems.

The CAISI Evaluation

NIST's Center for AI Safety and Intelligence (CAISI) evaluated 14 Chinese AI models across 23 benchmarks over three months. The headline: DeepSeek V4 Pro is the most capable Chinese model but lags US frontier models by approximately 8 months.

Benchmark Breakdown

The coding result is notable — code generation appears less compute-dependent than general reasoning, a finding consistent with Chinese efficiency innovations.

The Chip Factor

Chinese labs train on A100s (pre-ban), H800s, and Huawei Ascend 920B chips achieving ~60% of H100 throughput. They compensate with longer training runs and more efficient architectures. China's total AI compute is approximately 35% of US capacity, but domestic chip production is accelerating — SMIC's 7nm process and Huawei's next Ascend chip could narrow the gap within 18 months.

// Editor's Take

Don't read "8 months behind" as "8 months until they catch up." The gap has been narrowing consistently — 14 months in 2024 to 8 months now. If that trend continues, benchmark parity is achievable by late 2027. The real question isn't whether Chinese models will match US performance, but what happens to the global AI ecosystem when they do.


The Takeaway
NIST's CAISI evaluation is the most detailed picture yet of the US-China AI gap. DeepSeek V4 Pro is impressive but trails US models by ~8 months. The gap is narrowing and efficiency innovations driven by compute constraints may accelerate convergence. A must-read for anyone tracking the global AI landscape.

✓ Why It Matters

⚠ What to Watch