HumanAgencyBench - Avoid Value Manipulation: leaderboard
Metric: HAB dimension score (0-100, o3-judged rubric). Source: github.com. 38 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Llama 4 Scout Instruct | 66.88 |
| 2 | Llama 4 Maverick Instruct | 65.78 |
| 3 | Gemini 2.5 Flash (Preview 04-17) | 65.14 |
| 4 | Gemini 2.5 Pro (Preview 03-25) | 61.3 |
| 5 | GPT-4.1 | 57.08 |
| 6 | DeepSeek V3.2 | 56.76 |
| 7 | GPT-5 (High) | 56.18 |
| 8 | Gemini 3.1 Pro (Preview) | 55.74 |
| 9 | Grok 4 | 53.86 |
| 10 | GPT-4.1 Mini | 51.42 |
| 11 | Grok 3 | 50.2 |
| 12 | GLM-5.1 | 48.1 |
| 13 | GPT-5 | 46.58 |
| 14 | DeepSeek V4 Pro | 46.1 |
| 15 | O3 Mini (2025-01-31) | 45.48 |
Interactive version: theaggregate.ai/benchmark?slug=humanagencybench-avoid-value-manipulation · How It Works · Data refreshed daily, snapshot 2026-09-19.