VideoVIBE: leaderboard

Metric: Failure-diagnosis accuracy (%; unweighted macro-average of the twelve subdimension accuracies; zero-shot single-choice diagnosis from the interaction video and the page source code (Video+HTML), about 1.7K questions built from 6,338 verified failures of one-shot generated web applications recorded under human operation; free-form answers mapped deterministically to an option). Source: arxiv.org. Saturation forecast: Around 2028. 13 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash64.54
2Gemini 3.1 Pro (Preview)64.25
3Qwen 3.5 35B A3B63.9
4Qwen 3 VL 30B A3B (Thinking)60.96
5Qwen 3 VL 8B (Thinking)59.92
6Gemini 3.1 Flash Lite58.21
7Gemini 3 Flash57.09
8GLM-4.6V53.78
9Qwen 3 VL 30B A3B Instruct52.27
10MiMo-V2.549.61
11Qwen 3 VL 8B Instruct45.69

Interactive version: theaggregate.ai/benchmark?slug=videovibe · How It Works · Data refreshed daily, snapshot 2026-09-29.