VideoVIBE - Feature Execution: leaderboard

Metric: Failure-diagnosis accuracy (%; questions of the Feature Execution subdimension of Functional Executability; zero-shot single-choice diagnosis from the interaction video and the page source code (Video+HTML), about 1.7K questions built from 6,338 verified failures of one-shot generated web applications recorded under human operation; free-form answers mapped deterministically to an option). Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash90.37
2Gemini 3.1 Flash Lite89.84
3Qwen 3 VL 8B (Thinking)89.3
4Qwen 3 VL 30B A3B (Thinking)88.24
5Gemini 3.1 Pro (Preview)86.1
6Qwen 3 VL 30B A3B Instruct86.1
7Qwen 3.5 35B A3B83.96
8Gemini 3 Flash83.42
9Qwen 3 VL 8B Instruct79.68
10GLM-4.6V77.01
11MiMo-V2.573.26

Interactive version: theaggregate.ai/benchmark?slug=videovibe-feature-execution · How It Works · Data refreshed daily, snapshot 2026-09-29.