ScreenSpot-Pro — leaderboard

Professional GUI grounding benchmark requiring agents to identify precise screen locations in high-resolution development, creative, and scientific software.

Metric: Grounding score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 34 models tracked.

Top models

#ModelScore
1Claude Mythos Preview93
2Claude Mythos 590.7
3Claude Opus 4.889.5
4Claude Opus 4.787.6
5Qwen 3.7 Plus79
6Qwen 3.5 122B A10B70.4
7Qwen 3.5 27B70.3
8Qwen 3.5 35B A3B68.6
9Qwen 3.6 Plus68.2
10Gemini 3.1 Pro (Preview)68.1
11GPT-5.467.4
12Qwen 3.5 397B A17B65.6
13Qwen 3.5 9B65.2
14Qwen 3.5 4B60.3
15Qwen 3.5 2B54.5

Interactive version: theaggregate.ai/benchmark?slug=screenspot-pro · How the rankings work · Data refreshed daily, snapshot 2026-07-22.