OpenExempt — leaderboard

Diagnostic legal benchmark for reasoning over U.S. Bankruptcy Code exemption scenarios.

Metric: Mean F1 (self-reported). Source: benchmarklist.com. Status: saturation imminent. 13 models tracked.

Top models

#ModelScore
1O372.31
2Gemini 2.5 Pro71.71
3GPT-571.67
4Gemini 2.5 Flash65.76
5O4 Mini59.35
6Claude Sonnet 457.37
7GPT-4.147.77
8DeepSeek V347.67
9Llama 4 Maverick40.92
10Claude 3.5 Haiku36.13
11Llama 4 Scout31.91

Interactive version: theaggregate.ai/benchmark?slug=openexempt · How the rankings work · Data refreshed daily, snapshot 2026-07-22.