MT-Bench PL - Reasoning — leaderboard

Metric: Judge Score (0-10). Source: huggingface.co. 50 models tracked.

Top models

#ModelScore
1Phi-49.55
2Qwen 2.5 32B Instruct9.1
3Mistral Small 3.19
4aya-expanse-32B8.95
5Qwen 2 72B Instruct8.85
6Mistral Large 2 (Jul)8.7
7Gemma 3 27B (IT)8.4
8Bielik-11B-v2.3-Instruct8.35
9Mistral Small 37.9
10Mistral-Small-Instruct-24097.9
11Gemma 3 12B (IT)7.75
12Qwen 2.5 14B Instruct7.55
13Bielik-11B-v2.2-Instruct6.9
14Gemma 2 27B (IT)6.85
15aya-expanse-8B6.85

Interactive version: theaggregate.ai/benchmark?slug=mt-bench-pl-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.