MT-Bench PL - Reasoning: leaderboard

Metric: Judge Score (0-10). Source: huggingface.co. 50 models tracked.

Top models

#ModelScore
1Phi-49.55
2Qwen 2.5 32B Instruct9.1
3Mistral Small 3.19
4aya-expanse-32B8.95
5Qwen 2 72B Instruct8.85
6Mistral Large 2 (Jul)8.7
7Gemma 3 27B (IT)8.4
8Mistral Small 37.9
9Mistral-Small-Instruct-24097.9
10Gemma 3 12B (IT)7.75
11Qwen 2.5 14B Instruct7.55
12Bielik-11B-v2.2-Instruct6.9
13Gemma 2 27B (IT)6.85
14aya-expanse-8B6.85
15Mixtral 8x22B6.3

Interactive version: theaggregate.ai/benchmark?slug=mt-bench-pl-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.