ExploitBench v8-bench — leaderboard

Cybersecurity agent benchmark built from patched V8 exploitation tasks, grading progress from coverage and crash reproduction through exploit primitives.

Metric: Mean Capability (%). Source: exploitbench.ai. Status: saturation imminent. 9 models tracked.

Top models

#ModelScore
1Claude Mythos Preview78
2GPT-5.572
3Claude Opus 4.728
4Claude Sonnet 4.626
5Gemini 3.1 Pro (Preview)26
6GLM-5.118
7Kimi K2.618
8Claude Haiku 4.514
9MiniMax-M2.713

Interactive version: theaggregate.ai/benchmark?slug=exploitbench-v8-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.