MSVBench (Human Rating) - Visual Quality: leaderboard

Metric: Mean opinion score (1 to 5) for visual quality (clarity, lighting, stylistic unity, generative artifacts) (MSVBench human rating: 10 annotators with animation backgrounds rate each system's multi-shot storytelling videos generated from the MSVBench scripts (20 stories adapted from ViStoryBench, with character and shot reference images) on a 1-to-5 scale); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.

Top models

#ModelScoreOverall rank
1Veo3.14.29
2Sora24.22
3AniSora3.23.98
4LTXV-13B-0.9.83.9
5Wan2.2-I2V-A14B3.7
6Wan2.2-T2V-A14B3.22
7Wan2.2-TI2V-5B3.04
8Self-Forcing2.96
9HunyuanVideo-I2V2.71
10LongLive2.67

Interactive version: theaggregate.ai/benchmark?slug=msvbench-human-rating-visual-quality · How It Works · Data refreshed daily, snapshot 2026-10-11.