OSWorld 2.0 — leaderboard

Long-horizon computer-use benchmark with 108 end-to-end desktop workflows. Tests agents on realistic GUI, file, browser, and application tasks with binary completion as the primary metric.

Metric: Binary Accuracy (%). Source: osworld-v2.xlang.ai. Status: years away from saturation. 7 models tracked.

Top models

#ModelScore
1Claude Opus 4.820.6
2Claude Opus 4.718.2
3GPT-5.513
4Claude Sonnet 4.69.3
5Kimi K2.64.6
6MiniMax-M34.6
7Qwen 3.7 Plus2.8

Interactive version: theaggregate.ai/benchmark?slug=osworld-2-0 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.