Professional Reasoning Bench - Legal: leaderboard

Professional Reasoning Bench Legal evaluates frontier LLMs on complex legal reasoning tasks drawn from real-world legal practice and case analysis.

Metric: Score. Source: scale.com. Status: years away from saturation. 32 models tracked.

Top models

#ModelScore
1Muse Spark 1.157.05
2Claude Fable 552.56
3Muse Spark52.29
4Claude Opus 4.6 (Non-reasoning)52.27
5Claude Fable 5.1 (xHigh)51.6
6GPT-5.6 Sol (Max)50.5
7GPT-5 Pro49.89
8O3 Pro49.67
9GPT-5.1 (Thinking)49.33
10GPT-548.96
11O348.57
12GPT-5.2 Pro45.44
13GPT-5.4 (High)44.35
14Claude Opus 4.5 (20251101) (Thinking)44.21
15Gemini 3.1 Pro (Preview)44.02

Interactive version: theaggregate.ai/benchmark?slug=professional-reasoning-bench-legal · How It Works · Data refreshed daily, snapshot 2026-09-05.