World Agent - Relation: leaderboard

Metric: Maintenance relation score (%): share of authored causal, temporal and concurrency relations between realized events that hold, computed by program, mean per run over the nine-case common subset (27 runs per model, PhyAgentOS harness); higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 8 models tracked.

Top models

#ModelScore
1Qwen 3.8 Max (0902)73.8
2Qwen 3.8 Flash62.7
3DeepSeek V4.1 Flash58
4GLM-5.3 Flash57.4
5Kimi K354.6
6GPT-5.6 Luna54.3
7GLM-5.351.6
8Gemini 3.8 Flash51.2

Interactive version: theaggregate.ai/benchmark?slug=world-agent-relation · How It Works · Data refreshed daily, snapshot 2026-09-29.