VI-Bench (Video Prompt Inversion) - Easy (Semantic Grounding): leaderboard
Metric: Inversion Score (0-1; mean of the Prompt Score, GPT-4o (gpt-4o-2024-11-20) judging the recovered prompt against the original one on subject, action, scene, style and camera, and the Video Score, a Qwen3-VL-8B memory agent and Qwen3.5-9B judge agent comparing the video regenerated from the recovered prompt by the original generator (Wan2.2 or HunyuanVideo 1.5, fixed seed) with the reference, each a 1-5 rating normalized to 0-1; 300 single-shot videos controlling subject, action and scene). Source: arxiv.org. Saturation forecast: Around December 2026. 17 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Seed 2.0 Pro | 0.75 |
| 2 | Qwen 3 VL 8B | 0.73 |
| 3 | GPT-4o (2024-11-20) | 0.73 |
| 4 | Qwen 3.5 9B | 0.71 |
| 5 | Keye-VL-1.5-8B | 0.7 |
| 6 | Qwen 2.5 VL 7B | 0.7 |
| 7 | InternVL3-8B | 0.55 |
Interactive version: theaggregate.ai/benchmark?slug=vi-bench-video-prompt-inversion-easy-semantic-grounding · How It Works · Data refreshed daily, snapshot 2026-09-26.