Finance Agent v1.1 — leaderboard

Finance Agent v1.1 evaluates model capability on finance tasks from the linked upstream source with Score as the primary reported metric.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturated. 56 models tracked.

Top models

#ModelScore
1Claude Opus 4.764.4
2Claude Opus 4.7 (High)64.37
3Claude Sonnet 4.6 (Max)63.33
4GPT-5.4 Pro (xHigh)61.5
5Muse Spark60.59
6DeepSeek V4 Pro (Max)60.39
7Claude Opus 4.660.1
8Claude Opus 4.6 (Max)60.05
9GPT-5.560
10GPT-5.5 (xHigh)59.96
11Gemini 3.1 Pro (Preview) (High)59.72
12Gemini 3.1 Pro (Preview)59.7
13Claude Opus 4.5 (High)58.81
14GPT-5.2 (xHigh)58.53
15GLM-5.157.66

Interactive version: theaggregate.ai/benchmark?slug=finance-agent-v1-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.