Argo-Bench: leaderboard
Metric: Mean task score (0-100) over 210 enterprise data-science tasks on a simulated Oracle EBS warehouse (v1.1, graded against the simulator's hidden state on a private-seed world); one row per reasoning effort. Source: argo-bench.com. Saturation forecast: Around July 2028. 43 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 5.5 (xHigh) | 59.49 |
| 2 | Claude Opus 5.5 (High) | 55.55 |
| 3 | GPT-6 (xHigh) | 51.82 |
| 4 | Claude Sonnet 5.5 (xHigh) | 51.75 |
| 5 | Claude Opus 5.5 (Medium) | 49.64 |
| 6 | GPT-6.1 Sol (xHigh) | 49.53 |
| 7 | GPT-6 (High) | 47.55 |
| 8 | GPT-6 (Medium) | 45.11 |
| 9 | GPT-6.1 Sol (High) | 43.63 |
| 10 | Claude Sonnet 5.5 (High) | 40.33 |
| 11 | GPT-6 Sol (xHigh) | 36.75 |
| 12 | GPT-6.1 Sol (Medium) | 35.88 |
| 13 | GPT-6 (Low) | 33.32 |
| 14 | Claude Opus 5.5 (Low) | 32.56 |
| 15 | GPT-6 Sol (High) | 30.66 |
Interactive version: theaggregate.ai/benchmark?slug=argo-bench · How It Works · Data refreshed daily, snapshot 2026-10-01.