MV-Bench Multi-View Interfaces: leaderboard

Metric: Overall score (0-100): executability gate times the mean of static visual fidelity, data-binding correctness and interaction completeness, averaged per interface over the full-benchmark evaluation subset (72 base interfaces and 31 derived samples); each model is invoked as a coding agent through the Claude Agent SDK to build a React, TypeScript and D3 coordinated multi-view interface from a screenshot, dataset and interaction specification; single pass from one fixed prompt, no execution feedback; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 5 models tracked.

Top models

#ModelScore
1Kimi K2.530.9
2Qwen 3.5 Plus25.51
3Claude Sonnet 4.525.27
4GPT-5.425.25
5GLM-4.6V5.92

Interactive version: theaggregate.ai/benchmark?slug=mv-bench-multi-view-interfaces · How It Works · Data refreshed daily, snapshot 2026-09-29.