Odysseys: leaderboard

Metric: Perfect-rubric rate (%: share of tasks with every rubric item satisfied) over all 200 tasks of the Odysseys long-horizon multi-site web tasks run in an OSWorld Ubuntu desktop with Google Chrome, at most 100 steps, maximum reasoning effort, each rubric item judged from the trajectory screenshots and actions by gemini-3.1-flash-lite-preview; higher is better. Source: arxiv.org. Saturation forecast: Around October 2027. 8 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Max)44.5
2GPT-5.4 (xHigh)33.5
3Claude Sonnet 4.6 (Max)31
4Qwen 3.5 9B13.5
5Qwen 3.5 4B10.7
6GPT-5.4 Mini (xHigh)10.5
7Qwen 3.5 35B A3B6.5
8uitars-1.5-7B1

Interactive version: theaggregate.ai/benchmark?slug=odysseys · How It Works · Data refreshed daily, snapshot 2026-10-07.