ExploitBench v8-bench (AutoNudge): leaderboard

Metric: Capability coverage (%), all-runs view with reported seed/turn budgets; AutoNudge. Source: exploitbench.ai. Saturation forecast: Around January 2027. 10 models tracked.

Top models

#ModelScore
1Claude Mythos Preview78
2GPT-5.547
3Claude Opus 4.728
4Claude Sonnet 4.626
5Kimi K2.618
6GLM-5.118
7Gemini 3.1 Pro (Preview)16
8Claude Haiku 4.514
9MiniMax-M2.713

Interactive version: theaggregate.ai/benchmark?slug=exploitbench-v8-bench-autonudge · How It Works · Data refreshed daily, snapshot 2026-10-09.