ExploitGym: leaderboard

Real-world cybersecurity agent benchmark measuring whether AI agents can turn known software vulnerabilities into working, intended exploits across userspace, V8, and Linux kernel targets.

Metric: Successful Intended Exploits (#). Source: www.cybergym.io. Status: years away from saturation. 14 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol293
2Claude Mythos 5247
3Claude Opus 5191
4GLM-5.3130
5GPT-5.5129
6Claude Opus 4.8120
7GLM-5.279
8GPT-5.461
9Gemini 3.1 Pro (Preview)12
10Claude Opus 4.712
11Muse Spark 1.17
12GLM-5.14

Interactive version: theaggregate.ai/benchmark?slug=exploitgym · How It Works · Data refreshed daily, snapshot 2026-09-05.