MV-Bench Multi-View Interfaces (Repair): leaderboard

Metric: Overall score (0-100): executability gate times the mean of static visual fidelity, data-binding correctness and interaction completeness, averaged per interface over the full-benchmark evaluation subset (72 base interfaces and 31 derived samples); each model is invoked as a coding agent through the Claude Agent SDK to build a React, TypeScript and D3 coordinated multi-view interface from a screenshot, dataset and interaction specification; repair setting that feeds back the reference pipeline verification signals; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 5 models tracked.

Top models

#ModelScore
1Kimi K2.536.65
2Claude Sonnet 4.534.81
3Qwen 3.5 Plus30.56
4GPT-5.427.34
5GLM-4.6V11.98

Interactive version: theaggregate.ai/benchmark?slug=mv-bench-multi-view-interfaces-repair · How It Works · Data refreshed daily, snapshot 2026-09-29.