WBench - Physical Plausibility: leaderboard
Metric: Score (0-100): the mean of causal fidelity and visual plausibility, on the 158 navigation cases, multi-turn generation, each model under its own text, camera or action control; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 20 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Wan 2.7 | 71.8 |
| 2 | LingBot-World | 71.2 |
| 3 | Kling 3.0 | 69.3 |
| 4 | LongCat-Video | 68.9 |
| 5 | Seedance 1.5 | 68.3 |
| 6 | Cosmos 2.5 | 67.4 |
| 7 | HY-Video 1.5 | 67.3 |
| 8 | Fantasy-World | 66.8 |
| 9 | HY-World 1.5 | 66.3 |
| 10 | Genie 3 | 65.7 |
Interactive version: theaggregate.ai/benchmark?slug=wbench-physical-plausibility · How It Works · Data refreshed daily, snapshot 2026-10-07.