Harvey Legal Agent Benchmark: leaderboard

Legal-agent benchmark for completing realistic legal workflows with all-pass grading, including held-out Harvey tasks and public legal-agent task sets.

Metric: All-Pass Task Success (self-reported). Source: benchmarklist.com. Status: years away from saturation. 5 models tracked.

Top models

#ModelScore
1Claude Opus 523.58
2Claude Fable 5.1 (Max)19.09
3Claude Mythos 516.91
4Claude Mythos Preview13.4
5Claude Fable 513.3
6Claude Opus 4.810.4
7Claude Sonnet 58.92
8Claude Sonnet 4.68
9Claude Opus 4.77.1
10Claude Opus 4.64.2
11GPT-5.52.1
12Gemini 3.5 Flash0.8
13Gemini 3.1 Pro (Preview)0

Interactive version: theaggregate.ai/benchmark?slug=harvey-legal-agent-benchmark · How It Works · Data refreshed daily, snapshot 2026-09-05.