WorldReasonBench - Human-Centric: leaderboard
Metric: Process-aware reasoning score (0-100): Acc_QA^0.8 times s_dyn^0.2, where Acc_QA is the binary accuracy of structured questions about the generated video answered by Qwen3.5-27B (extended thinking, 4 fps) and judged against ground truth, and s_dyn the mean of its temporal and mechanism question accuracies, on the 9 Human-Centric cases on the 80-case shared evaluation set of WorldReasonBench (image-plus-text to video: from an initial frame and an action instruction, the generator must produce a video reaching the correct future world state; one generated clip per case); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Sora2 | 44.7 |
| 2 | Seedance2.0 | 35.9 |
| 3 | Veo3.1-Fast | 35.1 |
| 4 | Wan2.6 | 34.5 |
| 5 | WorldReasonBench Kling (checkpoint unspecified) | 32.5 |
| 6 | LongCat-Video | 22.8 |
| 7 | Cosmos-Predict2.5 | 22.2 |
| 8 | LTX2.3 | 19.3 |
| 9 | UniVideo | 15.8 |
| 10 | Wan2.2-14B | 14.5 |
Interactive version: theaggregate.ai/benchmark?slug=worldreasonbench-human-centric · How It Works · Data refreshed daily, snapshot 2026-10-07.