MT-Bench PL - Math — leaderboard

Metric: Judge Score (0-10). Source: huggingface.co. 50 models tracked.

Top models

#ModelScore
1Gemma 3 27B (IT)8.25
2Qwen 2.5 14B Instruct8.1
3Mistral Small 3.17.85
4Mistral Small 37.83
5Gemma 2 27B (IT)7.8
6Mistral Large 2 (Jul)7.8
7Phi-47.7
8Bielik-11B-v2.3-Instruct7.7
9Qwen 2.5 32B Instruct7.6
10Gemma 3 12B (IT)7.45
11Gemma 3 4B (IT)7.4
12Mistral-Small-Instruct-24097
13Mixtral 8x22B6.9
14GPT-3.5 Turbo6.85
15Mistral Nemo Instruct (2407)6.7

Interactive version: theaggregate.ai/benchmark?slug=mt-bench-pl-math · How the rankings work · Data refreshed daily, snapshot 2026-07-22.