CorpFin v2: leaderboard

A private benchmark evaluating understanding of long-context credit agreements.

Metric: Accuracy (%). Source: www.vals.ai. Status: years away from saturation. 134 models tracked.

Top models

#ModelScore
1Claude Opus 573.19
2Claude Fable 5 (Max)71.83
3Kimi K371.56
4Muse Spark 1.1 (xHigh)71.29
5Muse Spark 1.2 (xHigh)70.94
6Inkling Small69.62
7Inkling68.57
8Grok 4.368.53
9GPT-5.5 (xHigh)68.42
10Kimi K2.5 (Thinking)68.26
11MiniMax-M368.1
12Qwen 3 Max (2026-01-23)68.03
13Grok 4.5 (High)67.41
14Claude Opus 4.6 (Adaptive Reasoning, Max Effort)67.02
15Claude Sonnet 5 (Max)66.98

Interactive version: theaggregate.ai/benchmark?slug=corpfin-v2 · How It Works · Data refreshed daily, snapshot 2026-09-05.