PhyGround: leaderboard

Metric: Overall human rating (1-5): half the mean of the three general dimensions (semantic alignment, physical temporal validity, object persistence) plus half the per-law physics rating pooled over solid-body, fluid and optical laws, human ratings on a 1-5 scale from a quality-controlled pool of 352 annotators (after filtering 459), on 2,000 videos generated from 250 text-plus-first-frame prompts with explicit physical outcomes, each model at its released inference defaults; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1Wan2.2-27B-A14B3.28
2Veo-3.13.28
3OmniWeaving3.1
4Cosmos-Predict2.5-14B2.91
5LTX-2.3-22B2.69
6Wan2.2-TI2V-5B2.63
7Cosmos-Predict2.5-2B2.58
8LTX-2-19B2.56

Interactive version: theaggregate.ai/benchmark?slug=phyground · How It Works · Data refreshed daily, snapshot 2026-10-07.