WorldReasonBench - Human-Centric: leaderboard

Metric: Process-aware reasoning score (0-100): Acc_QA^0.8 times s_dyn^0.2, where Acc_QA is the binary accuracy of structured questions about the generated video answered by Qwen3.5-27B (extended thinking, 4 fps) and judged against ground truth, and s_dyn the mean of its temporal and mechanism question accuracies, on the 9 Human-Centric cases on the 80-case shared evaluation set of WorldReasonBench (image-plus-text to video: from an initial frame and an action instruction, the generator must produce a video reaching the correct future world state; one generated clip per case); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.

Top models

#ModelScore
1Sora244.7
2Seedance2.035.9
3Veo3.1-Fast35.1
4Wan2.634.5
5WorldReasonBench Kling (checkpoint unspecified)32.5
6LongCat-Video22.8
7Cosmos-Predict2.522.2
8LTX2.319.3
9UniVideo15.8
10Wan2.2-14B14.5

Interactive version: theaggregate.ai/benchmark?slug=worldreasonbench-human-centric · How It Works · Data refreshed daily, snapshot 2026-10-07.