Odysseys - Hard: leaderboard
Metric: Perfect-rubric rate (%: share of tasks with every rubric item satisfied) on the 109 hard tasks of the Odysseys long-horizon multi-site web tasks run in an OSWorld Ubuntu desktop with Google Chrome, at most 100 steps, maximum reasoning effort, each rubric item judged from the trajectory screenshots and actions by gemini-3.1-flash-lite-preview; higher is better. Source: arxiv.org. Saturation forecast: Around 2032. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 (Max) | 11 |
| 2 | Qwen 3.5 9B | 4.6 |
| 3 | Qwen 3.5 4B | 3.8 |
| 4 | GPT-5.4 (xHigh) | 3.7 |
| 5 | GPT-5.4 Mini (xHigh) | 1.8 |
| 6 | Claude Sonnet 4.6 (Max) | 1.8 |
| 7 | Qwen 3.5 35B A3B | 0 |
| 8 | uitars-1.5-7B | 0 |
Interactive version: theaggregate.ai/benchmark?slug=odysseys-hard · How It Works · Data refreshed daily, snapshot 2026-10-07.