The CAISI Evaluation
NIST's Center for AI Safety and Intelligence (CAISI) evaluated 14 Chinese AI models across 23 benchmarks over three months. The headline: DeepSeek V4 Pro is the most capable Chinese model but lags US frontier models by approximately 8 months.
Benchmark Breakdown
The coding result is notable — code generation appears less compute-dependent than general reasoning, a finding consistent with Chinese efficiency innovations.
- Coding (SWE-bench): DeepSeek 71.2% vs Claude Opus 4.7's 83.5% — smallest gap
- Mathematics (MATH-500): 89.1% vs 96.4% — strong but not frontier
- Reasoning (ARC): 94.7% vs 98.2% — competitive but trailing
- Multilingual: 91.3% vs 88.1% — DeepSeek wins on Chinese tasks
- Safety (TruthfulQA): 85.2% vs 92.8% — significant accuracy gap
The Chip Factor
Chinese labs train on A100s (pre-ban), H800s, and Huawei Ascend 920B chips achieving ~60% of H100 throughput. They compensate with longer training runs and more efficient architectures. China's total AI compute is approximately 35% of US capacity, but domestic chip production is accelerating — SMIC's 7nm process and Huawei's next Ascend chip could narrow the gap within 18 months.
Don't read "8 months behind" as "8 months until they catch up." The gap has been narrowing consistently — 14 months in 2024 to 8 months now. If that trend continues, benchmark parity is achievable by late 2027. The real question isn't whether Chinese models will match US performance, but what happens to the global AI ecosystem when they do.