Finance Agent v1.1 — leaderboard
Finance Agent v1.1 evaluates model capability on finance tasks from the linked upstream source with Score as the primary reported metric.
Metric: Score (self-reported). Source: benchmarklist.com. Status: saturated. 56 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.7 | 64.4 |
| 2 | Claude Opus 4.7 (High) | 64.37 |
| 3 | Claude Sonnet 4.6 (Max) | 63.33 |
| 4 | GPT-5.4 Pro (xHigh) | 61.5 |
| 5 | Muse Spark | 60.59 |
| 6 | DeepSeek V4 Pro (Max) | 60.39 |
| 7 | Claude Opus 4.6 | 60.1 |
| 8 | Claude Opus 4.6 (Max) | 60.05 |
| 9 | GPT-5.5 | 60 |
| 10 | GPT-5.5 (xHigh) | 59.96 |
| 11 | Gemini 3.1 Pro (Preview) (High) | 59.72 |
| 12 | Gemini 3.1 Pro (Preview) | 59.7 |
| 13 | Claude Opus 4.5 (High) | 58.81 |
| 14 | GPT-5.2 (xHigh) | 58.53 |
| 15 | GLM-5.1 | 57.66 |
Interactive version: theaggregate.ai/benchmark?slug=finance-agent-v1-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.