The Multilingual Math Gap
Researchers from the University of Lisbon and Instituto Superior Tecnico introduced MATH-PT, a math reasoning benchmark specifically designed for European and Brazilian Portuguese. The results reveal a troubling gap: frontier AI models perform 15-25% worse on math in Portuguese compared to English.
This isn't a translation issue — the benchmark was carefully created by native-speaking mathematicians, not translated from English. The performance gap reflects a real deficiency in how LLMs process mathematical reasoning in non-English languages.
Why This Matters
The implications extend beyond Portuguese:
- Over 6 billion people speak languages other than English
- AI-powered education tools in math are useless if they can't reason accurately in local languages
- The performance gap likely exists in hundreds of languages that lack benchmarks entirely
- Mathematical notation varies across cultures (comma vs period for decimals, different fraction representations)
The 15-25% math performance gap is a symptom of a deeper problem: almost all frontier model training is English-centric. We're building AI that's brilliant in English and mediocre everywhere else. If AI is going to be a global technology — not a rich-country luxury — we need multilingual training data, multilingual benchmarks, and multilingual evaluation teams. MATH-PT is a step in the right direction.