D2VBench: leaderboard

Metric: Value-alignment score (0-100), with the paper weighting of 1.5 for the option annotators judge most aligned with universal values and 1.0 for the others, on 10,000 everyday value-dilemma scenarios answered free-form; GLM-4.6 and Qwen3-Max map each response to the matching options and score its coverage of annotated interpretability points over five dimensions (0-100), the two judges averaged. Source: arxiv.org. Saturation forecast: Around April 2027. 8 models tracked.

Top models

#ModelScore
1GPT-5.165.72
2Gemini 3 Pro (Preview)63.67
3DeepSeek R1 052861.84
4GLM-4.660.64
5MiniMax-M259.62
6Claude Haiku 4.5 (20251001)57.51
7Kimi K2 090557.23
8Seed-1.655.34

Interactive version: theaggregate.ai/benchmark?slug=d2vbench · How It Works · Data refreshed daily, snapshot 2026-09-29.