Odysseys - Medium: leaderboard

Metric: Perfect-rubric rate (%: share of tasks with every rubric item satisfied) on the 46 medium tasks of the Odysseys long-horizon multi-site web tasks run in an OSWorld Ubuntu desktop with Google Chrome, at most 100 steps, maximum reasoning effort, each rubric item judged from the trajectory screenshots and actions by gemini-3.1-flash-lite-preview; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 8 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Max)71.7
2GPT-5.4 (xHigh)58.7
3Claude Sonnet 4.6 (Max)50
4GPT-5.4 Mini (xHigh)19.6
5Qwen 3.5 9B17.4
6Qwen 3.5 4B8.7
7Qwen 3.5 35B A3B4.3
8uitars-1.5-7B0

Interactive version: theaggregate.ai/benchmark?slug=odysseys-medium · How It Works · Data refreshed daily, snapshot 2026-10-07.