← Back to Blog
⚡ 10% OFF Get 10% off any Z.AI Coding Plan — covers GLM 5.1, GLM 5 Turbo, and GLM 4.7 → Claim discount
April 30, 2026 7 min read

MATH-PT: Why We Need Math Benchmarks in More Languages Than English

A new benchmark reveals that AI math reasoning drops significantly in non-English languages, exposing a critical gap in multilingual AI capability.

The Multilingual Math Gap

Researchers from the University of Lisbon and Instituto Superior Tecnico introduced MATH-PT, a math reasoning benchmark specifically designed for European and Brazilian Portuguese. The results reveal a troubling gap: frontier AI models perform 15-25% worse on math in Portuguese compared to English.

This isn't a translation issue — the benchmark was carefully created by native-speaking mathematicians, not translated from English. The performance gap reflects a real deficiency in how LLMs process mathematical reasoning in non-English languages.

Why This Matters

The implications extend beyond Portuguese:

// Editor's Take

The 15-25% math performance gap is a symptom of a deeper problem: almost all frontier model training is English-centric. We're building AI that's brilliant in English and mediocre everywhere else. If AI is going to be a global technology — not a rich-country luxury — we need multilingual training data, multilingual benchmarks, and multilingual evaluation teams. MATH-PT is a step in the right direction.


The Takeaway
MATH-PT exposes a significant gap in AI math reasoning for non-English languages. The 15-25% performance drop in Portuguese suggests similar or worse gaps in other languages. As AI becomes a global educational tool, multilingual capability isn't optional — it's essential.

✓ Why It Matters

⚠ What to Watch