WorldReasonBench: leaderboard
Metric: Process-aware reasoning score (0-100): Acc_QA^0.8 times s_dyn^0.2, where Acc_QA is the binary accuracy of structured questions about the generated video answered by Qwen3.5-27B (extended thinking, 4 fps) and judged against ground truth, and s_dyn the mean of its temporal and mechanism question accuracies, overall on the 80-case shared evaluation set of WorldReasonBench (image-plus-text to video: from an initial frame and an action instruction, the generator must produce a video reaching the correct future world state; one generated clip per case); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Seedance2.0 | 39.8 |
| 2 | Veo3.1-Fast | 35.3 |
| 3 | Sora2 | 34.3 |
| 4 | WorldReasonBench Kling (checkpoint unspecified) | 32.7 |
| 5 | Wan2.6 | 32.4 |
| 6 | HunyuanVideo-1.5 | 17.9 |
| 7 | Wan2.2-14B | 17.5 |
| 8 | LongCat-Video | 17.4 |
| 9 | Cosmos-Predict2.5 | 16.9 |
| 10 | LTX2.3 | 16.8 |
Interactive version: theaggregate.ai/benchmark?slug=worldreasonbench · How It Works · Data refreshed daily, snapshot 2026-10-07.