MT-Bench PL - Writing — leaderboard

Metric: Judge Score (0-10). Source: huggingface.co. 50 models tracked.

Top models

#ModelScore
1Gemma 3 27B (IT)9.7
2aya-expanse-32B9.6
3Bielik-11B-v2.3-Instruct9.5
4Bielik-11B-v2.1-Instruct9.5
5Mixtral 8x7B9.35
6Bielik-11B-v2.2-Instruct9.35
7Gemma 3 12B (IT)9.3
8Gemma 3 4B (IT)9.3
9aya-expanse-8B9.3
10Phi-49.25
11Mixtral 8x22B9.25
12Llama 3.1 405B Instruct9.2
13Mistral Small 3.19.15
14Llama 3.1 70B Instruct9.1
15GPT-3.5 Turbo9.1

Interactive version: theaggregate.ai/benchmark?slug=mt-bench-pl-writing · How the rankings work · Data refreshed daily, snapshot 2026-07-22.