MT-Bench PL - Writing: leaderboard

Metric: Judge Score (0-10). Source: huggingface.co. 50 models tracked.

Top models

#ModelScore
1Gemma 3 27B (IT)9.7
2aya-expanse-32B9.6
3Bielik-11B-v2.1-Instruct9.5
4Mixtral 8x7B9.35
5Bielik-11B-v2.2-Instruct9.35
6Gemma 3 12B (IT)9.3
7Gemma 3 4B (IT)9.3
8aya-expanse-8B9.3
9Phi-49.25
10Mixtral 8x22B9.25
11Llama 3.1 405B Instruct9.2
12Mistral Small 3.19.15
13Llama 3.1 70B Instruct9.1
14GPT-3.5 Turbo9.1
15Mistral-Small-Instruct-24098.8

Interactive version: theaggregate.ai/benchmark?slug=mt-bench-pl-writing · How It Works · Data refreshed daily, snapshot 2026-09-05.