Argo-Bench - Solved: leaderboard

Metric: Share of the 210 tasks solved (%), a task counting as solved at a score of 95 or more (v1.1); one row per reasoning effort. Source: argo-bench.com. Saturation forecast: Around 2030. 43 models tracked.

Top models

#ModelScore
1Claude Opus 5.5 (xHigh)34.76
2Claude Opus 5.5 (High)31.9
3Claude Sonnet 5.5 (xHigh)28.57
4Claude Opus 5.5 (Medium)28.1
5GPT-6 (xHigh)27.62
6GPT-6 (High)24.76
7GPT-6.1 Sol (xHigh)24.76
8GPT-6.1 Sol (High)23.33
9Claude Sonnet 5.5 (High)22.86
10GPT-6 (Medium)22.38
11GPT-6.1 Sol (Medium)19.52
12Claude Opus 5.5 (Low)19.52
13DeepSeek V4.1 Flash (Max)17.62
14GPT-6 Sol (xHigh)17.62
15GPT-6 (Low)17.14

Interactive version: theaggregate.ai/benchmark?slug=argo-bench-solved · How It Works · Data refreshed daily, snapshot 2026-10-01.