Math-PT (pt-BR MC) - Level 4: leaderboard

Metric: Accuracy (%) on the 270 level-4 multiple-choice questions in Brazilian Portuguese (OBMEP and OMIF); native-language zero-shot prompt asking for a boxed final answer, figures given as LaTeX source; the letter in the final boxed answer must match the key; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.

Top models

#ModelScore
1GPT-593.33
2Qwen 3 235B A22B91.48
3DeepSeek V3.189.26
4Qwen 3 14B88.15
5Gemini 2.5 Flash85.93
6Qwen 3 32B85.19
7Qwen 3 30B A3B84.44
8Claude Haiku 4.581.48
9Qwen 3 8B81.48
10Gemma 3 27B (IT)64.81
11Llama 3.3 70B Instruct32.96

Interactive version: theaggregate.ai/benchmark?slug=math-pt-pt-br-mc-level-4 · How It Works · Data refreshed daily, snapshot 2026-10-07.