SpatialWorld - Work: leaderboard

Metric: Task success rate (%) on the 59 work tasks; task success rate (%) judged by terminal-state verifiers, vision-only egocentric observation and a unified text action interface, temperature 1.0, 30-turn history, step budget twice the human golden action count plus 10; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 15 models tracked.

Top models

#ModelScore
1GPT-516.9
2Qwen 3.5 397B A17B16.9
3Gemini 2.5 Pro11.9
4Gemini 3.1 Pro (Preview)10.2
5Gemini 3 Flash10.2
6Kimi K2.58.5
7Qwen 3 VL 235B A22B Instruct8.5
8Qwen 2.5 VL 72B Instruct8.5
9Qwen 3 VL 235B A22B (Thinking)8.5
10Seed 2.0 Lite6.8
11GPT-5.45.1
12GLM-4.6V5.1
13GLM-4.5V3.4

Interactive version: theaggregate.ai/benchmark?slug=spatialworld-work · How It Works · Data refreshed daily, snapshot 2026-09-29.