WorldReasonBench - World Knowledge: leaderboard
Metric: Process-aware reasoning score (0-100): Acc_QA^0.8 times s_dyn^0.2, where Acc_QA is the binary accuracy of structured questions about the generated video answered by Qwen3.5-27B (extended thinking, 4 fps) and judged against ground truth, and s_dyn the mean of its temporal and mechanism question accuracies, on the 21 World Knowledge cases on the 80-case shared evaluation set of WorldReasonBench (image-plus-text to video: from an initial frame and an action instruction, the generator must produce a video reaching the correct future world state; one generated clip per case); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Veo3.1-Fast | 55 |
| 2 | Seedance2.0 | 43.2 |
| 3 | WorldReasonBench Kling (checkpoint unspecified) | 42.2 |
| 4 | Sora2 | 36.9 |
| 5 | Wan2.6 | 35.2 |
| 6 | Wan2.2-14B | 22.9 |
| 7 | HunyuanVideo-1.5 | 21.6 |
| 8 | LTX2.3 | 15.6 |
| 9 | Cosmos-Predict2.5 | 15.2 |
| 10 | UniVideo | 13.8 |
Interactive version: theaggregate.ai/benchmark?slug=worldreasonbench-world-knowledge · How It Works · Data refreshed daily, snapshot 2026-10-07.