Agent-ValueBench — leaderboard

Agent value-alignment benchmark with executable environments and value-conflict tasks, testing whether autonomous agents express stable values across domains, harnesses, and trajectories.

Metric: Authority (self-reported). Source: benchmarklist.com. 14 models tracked.

Top models

#ModelScore
1Qwen 3 30B A3B7.8
2Grok 4.207.8
3Llama 3.3 70B Instruct7.5
4Gemini 3.1 Pro (Preview)6.7
5MiniMax-M2.75.6
6GPT-5.4 Mini5.5
7Claude Sonnet 4.65
8Claude Haiku 4.54.7
9Qwen 3.5 397B A17B4.7
10GPT-5.44.6
11GLM-5.14.4
12DeepSeek V3.23.4
13Gemini 3 Flash (Preview)2.1

Interactive version: theaggregate.ai/benchmark?slug=agent-valuebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.