PAWBench (World Modeling) - Calibration: leaderboard
Metric: Conditional total-variation distance between the produced and reference outcome distributions for video generators, times 100, on a 0 to 100 scale (0-100); lower is better; PAWBench measures whether a system reproduces the reference distribution over physically possible outcomes across 50 rollouts on 25 scenarios per physical track, averaged over scenes that pass the readability gate. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Cosmos 3 Super I2V | 20.5 |
| 2 | MiniMax H3 | 24.2 |
| 3 | Wan2.7 | 26.3 |
| 4 | Wan2.2 | 26.3 |
| 5 | LTX-2.3 | 30.1 |
| 6 | LTX-2.5 | 30.2 |
| 7 | Seedance 2 | 30.5 |
| 8 | Kling 3 Std. | 34.9 |
| 9 | Veo3.1 Fast | 35.4 |
| 10 | LingBot-Video-MoE | 41.8 |
Interactive version: theaggregate.ai/benchmark?slug=pawbench-world-modeling-calibration · How It Works · Data refreshed daily, snapshot 2026-09-26.