OSWorld — leaderboard

Computer-use agent benchmark with 369 real-world tasks across Ubuntu, Windows, and macOS. Tests ability to interact with GUIs, terminals, and applications via screenshots and actions.

Metric: Success Rate (%). Source: os-world.github.io. Status: saturation imminent. 63 models tracked.

Top models

#ModelScore
1Muse Spark 1.180.67
2MiniMax-M375.19
3Qwen 3.7 Plus73.3
4Kimi K2.673.06
5Median Human72.36
6Claude Sonnet 4.672.11
7Kimi K2.563.3
8Claude Sonnet 4.562.88
9Seed 1.861.87
10Claude Sonnet 4 (20250514)43.9
11Claude 3.7 Sonnet (20250219)35.8
12uitars-1.5-7B29.6
13O323
14Kimi-VL-A3B10.3
15Qwen 2.5 VL 32B Instruct3.88

Interactive version: theaggregate.ai/benchmark?slug=osworld · How the rankings work · Data refreshed daily, snapshot 2026-07-22.