VI-Bench (Video Prompt Inversion) - Easy (Semantic Grounding): leaderboard

Metric: Inversion Score (0-1; mean of the Prompt Score, GPT-4o (gpt-4o-2024-11-20) judging the recovered prompt against the original one on subject, action, scene, style and camera, and the Video Score, a Qwen3-VL-8B memory agent and Qwen3.5-9B judge agent comparing the video regenerated from the recovered prompt by the original generator (Wan2.2 or HunyuanVideo 1.5, fixed seed) with the reference, each a 1-5 rating normalized to 0-1; 300 single-shot videos controlling subject, action and scene). Source: arxiv.org. Saturation forecast: Around December 2026. 17 models tracked.

Top models

#ModelScore
1Seed 2.0 Pro0.75
2Qwen 3 VL 8B0.73
3GPT-4o (2024-11-20)0.73
4Qwen 3.5 9B0.71
5Keye-VL-1.5-8B0.7
6Qwen 2.5 VL 7B0.7
7InternVL3-8B0.55

Interactive version: theaggregate.ai/benchmark?slug=vi-bench-video-prompt-inversion-easy-semantic-grounding · How It Works · Data refreshed daily, snapshot 2026-09-26.