MT-Bench PL - Math: leaderboard

Metric: Judge Score (0-10). Source: huggingface.co. 50 models tracked.

Top models

#ModelScore
1Gemma 3 27B (IT)8.25
2Qwen 2.5 14B Instruct8.1
3Mistral Small 3.17.85
4Mistral Small 37.83
5Gemma 2 27B (IT)7.8
6Mistral Large 2 (Jul)7.8
7Phi-47.7
8Qwen 2.5 32B Instruct7.6
9Gemma 3 12B (IT)7.45
10Gemma 3 4B (IT)7.4
11Mistral-Small-Instruct-24097
12Mixtral 8x22B6.9
13GPT-3.5 Turbo6.85
14Mistral Nemo Instruct (2407)6.7
15aya-expanse-32B6.6

Interactive version: theaggregate.ai/benchmark?slug=mt-bench-pl-math · How It Works · Data refreshed daily, snapshot 2026-09-05.