MV-Bench Multi-View Interfaces: leaderboard
Metric: Overall score (0-100): executability gate times the mean of static visual fidelity, data-binding correctness and interaction completeness, averaged per interface over the full-benchmark evaluation subset (72 base interfaces and 31 derived samples); each model is invoked as a coding agent through the Claude Agent SDK to build a React, TypeScript and D3 coordinated multi-view interface from a screenshot, dataset and interaction specification; single pass from one fixed prompt, no execution feedback; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Kimi K2.5 | 30.9 |
| 2 | Qwen 3.5 Plus | 25.51 |
| 3 | Claude Sonnet 4.5 | 25.27 |
| 4 | GPT-5.4 | 25.25 |
| 5 | GLM-4.6V | 5.92 |
Interactive version: theaggregate.ai/benchmark?slug=mv-bench-multi-view-interfaces · How It Works · Data refreshed daily, snapshot 2026-09-29.