MSVBench (Human Rating): leaderboard

Metric: Mean opinion score (1 to 5) averaged over visual quality, story-video alignment, video consistency and motion quality (MSVBench human rating: 10 annotators with animation backgrounds rate each system's multi-shot storytelling videos generated from the MSVBench scripts (20 stories adapted from ViStoryBench, with character and shot reference images) on a 1-to-5 scale); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.

Top models

#ModelScoreOverall rank
1Veo3.14.11
2Sora24.06
3AniSora3.23.43
4Wan2.2-I2V-A14B3.36
5LTXV-13B-0.9.83.36
6Wan2.2-T2V-A14B3.08
7Wan2.2-TI2V-5B2.92
8HunyuanVideo-I2V2.76
9Self-Forcing2.69
10CogVideoX1.5-5B-I2V2.39

Interactive version: theaggregate.ai/benchmark?slug=msvbench-human-rating · How It Works · Data refreshed daily, snapshot 2026-10-11.