DV-World - DV-Inter: leaderboard

Metric: DV-Inter score (%): per-task MLLM rubric score (interaction, accuracy, aesthetics) times the interaction success rate, on ambiguous visualization requests resolved with a GPT-5-mini user simulator driven by expert hidden intents, run in the DV-World-Agent framework (data manipulation, multimodal perception and proactive interaction tools); expert rubrics applied by a Gemini-2.5-Flash judge where rubric-scored, means over evaluation trials; higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 10 models tracked.

Top models

#ModelScore
1Grok 440.43
2DeepSeek V3.237.94
3GPT-5.235.09
4Gemini 3 Pro (Preview)34.43
5Gemini 2.5 Pro31.34
6GLM-4.729.57
7Kimi K2 (Thinking)27.39
8GPT-4.125.68
9Qwen 3 235B A22B20.9
10Qwen 3 8B18.09

Interactive version: theaggregate.ai/benchmark?slug=dv-world-dv-inter · How It Works · Data refreshed daily, snapshot 2026-10-07.