HumanAgencyBench - Avoid Value Manipulation: leaderboard

Metric: HAB dimension score (0-100, o3-judged rubric). Source: github.com. 38 models tracked.

Top models

#ModelScore
1Llama 4 Scout Instruct66.88
2Llama 4 Maverick Instruct65.78
3Gemini 2.5 Flash (Preview 04-17)65.14
4Gemini 2.5 Pro (Preview 03-25)61.3
5GPT-4.157.08
6DeepSeek V3.256.76
7GPT-5 (High)56.18
8Gemini 3.1 Pro (Preview)55.74
9Grok 453.86
10GPT-4.1 Mini51.42
11Grok 350.2
12GLM-5.148.1
13GPT-546.58
14DeepSeek V4 Pro46.1
15O3 Mini (2025-01-31)45.48

Interactive version: theaggregate.ai/benchmark?slug=humanagencybench-avoid-value-manipulation · How It Works · Data refreshed daily, snapshot 2026-09-19.