PLawBench — leaderboard

Polish legal benchmark for practical legal-work evaluation, with legal questions and detailed rubric-based assessment.

Metric: Overall (self-reported). Source: benchmarklist.com. Status: saturation imminent. 24 models tracked.

Top models

#ModelScore
1GPT-5.269.67
2GPT-567.76
3Claude Opus 4.566.47
4Gemini 3 Pro (Preview)66.35
5Claude Sonnet 4.565.88
6Gemini 2.5 Pro64.05
7Qwen 3 235B A22B 2507 Instruct63.08
8Gemini 2.5 Flash62.07
9GLM-4.660.49
10DeepSeek V3.257.97
11Qwen 3 30B A3B 2507 Instruct55.73
12Claude Sonnet 453.55
13Grok 4.1 Fast52.67
14ERNIE 5.0 Thinking Preview51.39
15Qwen 3 4B 2507 Instruct50.53

Interactive version: theaggregate.ai/benchmark?slug=plawbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.