PAWBench (World Modeling) - Coverage: leaderboard
Metric: Coverage (%; PAWBench measures whether a system reproduces the reference distribution over physically possible outcomes across 50 rollouts on 25 scenarios per physical track with a generated video; valid-support recovery: share of reference-supported outcomes the model produces (%, higher is better; averaged over passing scenes)). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | LTX-2.3 | 71.7 |
| 2 | Wan2.2 | 63.4 |
| 3 | LingBot-Video-MoE | 58.8 |
| 4 | LTX-2.5 | 57.4 |
| 5 | Cosmos 3 Super I2V | 55.2 |
| 6 | Kling 3 Std. | 52.8 |
| 7 | Seedance 2 | 50.9 |
| 8 | Wan2.7 | 50 |
| 9 | MiniMax H3 | 48.7 |
| 10 | HappyHorse | 47.1 |
Interactive version: theaggregate.ai/benchmark?slug=pawbench-world-modeling-coverage · How It Works · Data refreshed daily, snapshot 2026-09-26.