MSVBench (Human Rating): leaderboard
Metric: Mean opinion score (1 to 5) averaged over visual quality, story-video alignment, video consistency and motion quality (MSVBench human rating: 10 annotators with animation backgrounds rate each system's multi-shot storytelling videos generated from the MSVBench scripts (20 stories adapted from ViStoryBench, with character and shot reference images) on a 1-to-5 scale); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Veo3.1 | 4.11 | |
| 2 | Sora2 | 4.06 | |
| 3 | AniSora3.2 | 3.43 | |
| 4 | Wan2.2-I2V-A14B | 3.36 | |
| 5 | LTXV-13B-0.9.8 | 3.36 | |
| 6 | Wan2.2-T2V-A14B | 3.08 | |
| 7 | Wan2.2-TI2V-5B | 2.92 | |
| 8 | HunyuanVideo-I2V | 2.76 | |
| 9 | Self-Forcing | 2.69 | |
| 10 | CogVideoX1.5-5B-I2V | 2.39 |
Interactive version: theaggregate.ai/benchmark?slug=msvbench-human-rating · How It Works · Data refreshed daily, snapshot 2026-10-11.