WorldReasonBench: leaderboard

Metric: Process-aware reasoning score (0-100): Acc_QA^0.8 times s_dyn^0.2, where Acc_QA is the binary accuracy of structured questions about the generated video answered by Qwen3.5-27B (extended thinking, 4 fps) and judged against ground truth, and s_dyn the mean of its temporal and mechanism question accuracies, overall on the 80-case shared evaluation set of WorldReasonBench (image-plus-text to video: from an initial frame and an action instruction, the generator must produce a video reaching the correct future world state; one generated clip per case); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.

Top models

#ModelScore
1Seedance2.039.8
2Veo3.1-Fast35.3
3Sora234.3
4WorldReasonBench Kling (checkpoint unspecified)32.7
5Wan2.632.4
6HunyuanVideo-1.517.9
7Wan2.2-14B17.5
8LongCat-Video17.4
9Cosmos-Predict2.516.9
10LTX2.316.8

Interactive version: theaggregate.ai/benchmark?slug=worldreasonbench · How It Works · Data refreshed daily, snapshot 2026-10-07.