DV-World - DV-Inter: leaderboard
Metric: DV-Inter score (%): per-task MLLM rubric score (interaction, accuracy, aesthetics) times the interaction success rate, on ambiguous visualization requests resolved with a GPT-5-mini user simulator driven by expert hidden intents, run in the DV-World-Agent framework (data manipulation, multimodal perception and proactive interaction tools); expert rubrics applied by a Gemini-2.5-Flash judge where rubric-scored, means over evaluation trials; higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Grok 4 | 40.43 |
| 2 | DeepSeek V3.2 | 37.94 |
| 3 | GPT-5.2 | 35.09 |
| 4 | Gemini 3 Pro (Preview) | 34.43 |
| 5 | Gemini 2.5 Pro | 31.34 |
| 6 | GLM-4.7 | 29.57 |
| 7 | Kimi K2 (Thinking) | 27.39 |
| 8 | GPT-4.1 | 25.68 |
| 9 | Qwen 3 235B A22B | 20.9 |
| 10 | Qwen 3 8B | 18.09 |
Interactive version: theaggregate.ai/benchmark?slug=dv-world-dv-inter · How It Works · Data refreshed daily, snapshot 2026-10-07.