Argo-Bench - Solved: leaderboard
Metric: Share of the 210 tasks solved (%), a task counting as solved at a score of 95 or more (v1.1); one row per reasoning effort. Source: argo-bench.com. Saturation forecast: Around 2030. 43 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 5.5 (xHigh) | 34.76 |
| 2 | Claude Opus 5.5 (High) | 31.9 |
| 3 | Claude Sonnet 5.5 (xHigh) | 28.57 |
| 4 | Claude Opus 5.5 (Medium) | 28.1 |
| 5 | GPT-6 (xHigh) | 27.62 |
| 6 | GPT-6 (High) | 24.76 |
| 7 | GPT-6.1 Sol (xHigh) | 24.76 |
| 8 | GPT-6.1 Sol (High) | 23.33 |
| 9 | Claude Sonnet 5.5 (High) | 22.86 |
| 10 | GPT-6 (Medium) | 22.38 |
| 11 | GPT-6.1 Sol (Medium) | 19.52 |
| 12 | Claude Opus 5.5 (Low) | 19.52 |
| 13 | DeepSeek V4.1 Flash (Max) | 17.62 |
| 14 | GPT-6 Sol (xHigh) | 17.62 |
| 15 | GPT-6 (Low) | 17.14 |
Interactive version: theaggregate.ai/benchmark?slug=argo-bench-solved · How It Works · Data refreshed daily, snapshot 2026-10-01.