VI-Bench (Video Prompt Inversion) - Hard (Multi-Shot): leaderboard

Metric: Inversion Score (0-1; mean of the Prompt Score, GPT-4o (gpt-4o-2024-11-20) judging the recovered prompt against the original one on subject, action, scene, style and camera, and the Video Score, a Qwen3-VL-8B memory agent and Qwen3.5-9B judge agent comparing the video regenerated from the recovered prompt by the original generator (Wan2.2 or HunyuanVideo 1.5, fixed seed) with the reference, each a 1-5 rating normalized to 0-1; 300 videos of 2-4 shots, one shot-level prompt per shot, replayed shot by shot). Source: arxiv.org. Saturation forecast: Around May 2027. 17 models tracked.

Top models

#ModelScore
1Seed 2.0 Pro0.53
2GPT-4o (2024-11-20)0.44
3Qwen 2.5 VL 7B0.38
4Qwen 3 VL 8B0.37
5Qwen 3.5 9B0.3
6InternVL3-8B0.24
7Keye-VL-1.5-8B0.22

Interactive version: theaggregate.ai/benchmark?slug=vi-bench-video-prompt-inversion-hard-multi-shot · How It Works · Data refreshed daily, snapshot 2026-09-26.