WBench - Physical Plausibility: leaderboard

Metric: Score (0-100): the mean of causal fidelity and visual plausibility, on the 158 navigation cases, multi-turn generation, each model under its own text, camera or action control; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 20 models tracked.

Top models

#ModelScore
1Wan 2.771.8
2LingBot-World71.2
3Kling 3.069.3
4LongCat-Video68.9
5Seedance 1.568.3
6Cosmos 2.567.4
7HY-Video 1.567.3
8Fantasy-World66.8
9HY-World 1.566.3
10Genie 365.7

Interactive version: theaggregate.ai/benchmark?slug=wbench-physical-plausibility · How It Works · Data refreshed daily, snapshot 2026-10-07.