Harvey Legal Agent Benchmark — leaderboard

Legal-agent benchmark for completing realistic legal workflows with all-pass grading, including held-out Harvey tasks and public legal-agent task sets.

Metric: All-Pass Task Success (self-reported). Source: benchmarklist.com. Status: saturation imminent. 11 models tracked.

Top models

#ModelScore
1Claude Mythos 516.91
2Claude Mythos Preview13.4
3Claude Fable 513.3
4Claude Opus 4.810.4
5Claude Sonnet 58.92
6Claude Opus 4.77.1
7Claude Sonnet 4.65.4
8Claude Opus 4.64.2
9GPT-5.52.1
10Gemini 3.5 Flash0.8
11Gemini 3.1 Pro (Preview)0

Interactive version: theaggregate.ai/benchmark?slug=harvey-legal-agent-benchmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.