OSWorld-Verified — leaderboard

OSWorld-Verified evaluates model capability on agentic tasks from the linked upstream source with Score as the primary reported metric.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 20 models tracked.

Top models

#ModelScore
1Claude Fable 5 (Max)85
2Claude Opus 4.883.4
3Claude Opus 4.8 (Max)83.4
4Claude Opus 4.782.8
5Claude Sonnet 581.2
6Claude Mythos Preview79.6
7GPT-5.578.7
8Claude Sonnet 4.678.5
9Gemini 3.5 Flash78.4
10Gemini 3.1 Pro (Preview)76.2
11GPT-5.475
12Qwen 3.7 Plus73.3
13Claude Opus 4.6 (Max)72.7
14MiniMax-M370.06
15Qwen 3.6 Plus62.5

Interactive version: theaggregate.ai/benchmark?slug=osworld-verified · How the rankings work · Data refreshed daily, snapshot 2026-07-22.