CTFBench — leaderboard

CTFBench: Measures model robustness, truthfulness, calibration, bias, harmfulness, jailbreak resistance, or alignment-relevant behavior.

Metric: Vulnerability Detection Rate (self-reported). Source: benchmarklist.com. Status: saturation imminent. 27 models tracked.

Top models

#ModelScore
1GPT-5.5100
2Gemini 3.1 Pro (Preview)98.4
3Claude Opus 4.688.9
4Kimi K2.682.5
5Claude Opus 4.773
6GPT-571.4
7Claude Opus 4.571.4
8DeepSeek V4 Pro64.8
9Gemini 2.5 Pro57.1
10GPT-5.457.1
11MiMo-V2.5-Pro55.6
12GPT-5.252.4
13Grok 352.4
14MiniMax-M2.750.8
15DeepSeek R142.9

Interactive version: theaggregate.ai/benchmark?slug=ctfbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.