Reward Hacking Benchmark (RHB): leaderboard

Agentic reward-hacking benchmark with multi-step tool-use tasks containing naturalistic shortcut opportunities such as skipping verification, exploiting metadata, or tampering with evaluation-relevant functions.

Metric: Exploit rate (%) (self-reported). Source: benchmarklist.com. Status: saturation imminent. 13 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.50
2Claude Opus 4.50
3DeepSeek V30.6
4GPT-4o0.9
5Claude 3.7 Sonnet3.9
6Gemini 2.5 Pro (Preview)4.6
7O16.8
8O3 Mini7.1
9O4 Mini8.4
10O311.8

Interactive version: theaggregate.ai/benchmark?slug=reward-hacking-benchmark-rhb · How It Works · Data refreshed daily, snapshot 2026-09-05.