PhyGround - Fluid: leaderboard
Metric: Human rating (1-5) of fluid laws (liquid-solid interaction, flow dynamics, conservation), pooled over (video, law) units, human ratings on a 1-5 scale from a quality-controlled pool of 352 annotators (after filtering 459), on 2,000 videos generated from 250 text-plus-first-frame prompts with explicit physical outcomes, each model at its released inference defaults; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Veo-3.1 | 3.65 |
| 2 | OmniWeaving | 3.26 |
| 3 | Wan2.2-27B-A14B | 3.18 |
| 4 | LTX-2.3-22B | 3.07 |
| 5 | LTX-2-19B | 3.04 |
| 6 | Cosmos-Predict2.5-14B | 2.98 |
| 7 | Cosmos-Predict2.5-2B | 2.77 |
| 8 | Wan2.2-TI2V-5B | 2.71 |
Interactive version: theaggregate.ai/benchmark?slug=phyground-fluid · How It Works · Data refreshed daily, snapshot 2026-10-07.