World Agent: leaderboard

Metric: Maintenance completion score (%): share of a world's contracted events, structures and probes realized, mean per run over the nine-case common subset (27 runs per model, easy/hard/extreme tiers, PhyAgentOS harness, isolated LLM judge plus programmatic audit); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 8 models tracked.

Top models

#ModelScore
1Qwen 3.8 Max (0902)81.3
2GLM-5.3 Flash75.2
3DeepSeek V4.1 Flash74.4
4Qwen 3.8 Flash73.5
5GPT-5.6 Luna73.4
6Kimi K370.9
7Gemini 3.8 Flash70.4
8GLM-5.369.2

Interactive version: theaggregate.ai/benchmark?slug=world-agent · How It Works · Data refreshed daily, snapshot 2026-09-29.