ExploitGym — leaderboard

Real-world cybersecurity agent benchmark measuring whether AI agents can turn known software vulnerabilities into working, intended exploits across userspace, V8, and Linux kernel targets.

Metric: Successful Intended Exploits (#). Source: www.cybergym.io. Status: saturation imminent. 9 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol293
2GPT-5.5129
3GPT-5.461
4Gemini 3.1 Pro (Preview)12
5Claude Opus 4.712
6Muse Spark 1.17
7GLM-5.14

Interactive version: theaggregate.ai/benchmark?slug=exploitgym · How the rankings work · Data refreshed daily, snapshot 2026-07-22.