SpatialWorld - Social Collaboration: leaderboard

Metric: Task success rate (%) on the 46 multi-agent social collaboration tasks (Multi-AI2THOR and Multi-ProcTHOR); task success rate (%) judged by terminal-state verifiers, vision-only egocentric observation and a unified text action interface, temperature 1.0, 30-turn history, step budget twice the human golden action count plus 10; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 15 models tracked.

Top models

#ModelScore
1GPT-534.8
2Qwen 3.5 397B A17B19.6
3Kimi K2.517.4
4GLM-4.5V13
5Seed 2.0 Lite13
6Gemini 2.5 Pro10.9
7Qwen 3 VL 235B A22B Instruct10.9
8Qwen 3 VL 235B A22B (Thinking)10.9
9Gemini 3.1 Pro (Preview)8.7
10GPT-5.46.5
11Gemini 3 Flash4.3
12Qwen 2.5 VL 72B Instruct2.2
13GLM-4.6V0

Interactive version: theaggregate.ai/benchmark?slug=spatialworld-social-collaboration · How It Works · Data refreshed daily, snapshot 2026-09-29.