MSVBench (Human Rating) - Visual Quality: leaderboard
Metric: Mean opinion score (1 to 5) for visual quality (clarity, lighting, stylistic unity, generative artifacts) (MSVBench human rating: 10 annotators with animation backgrounds rate each system's multi-shot storytelling videos generated from the MSVBench scripts (20 stories adapted from ViStoryBench, with character and shot reference images) on a 1-to-5 scale); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Veo3.1 | 4.29 | |
| 2 | Sora2 | 4.22 | |
| 3 | AniSora3.2 | 3.98 | |
| 4 | LTXV-13B-0.9.8 | 3.9 | |
| 5 | Wan2.2-I2V-A14B | 3.7 | |
| 6 | Wan2.2-T2V-A14B | 3.22 | |
| 7 | Wan2.2-TI2V-5B | 3.04 | |
| 8 | Self-Forcing | 2.96 | |
| 9 | HunyuanVideo-I2V | 2.71 | |
| 10 | LongLive | 2.67 |
Interactive version: theaggregate.ai/benchmark?slug=msvbench-human-rating-visual-quality · How It Works · Data refreshed daily, snapshot 2026-10-11.