SecRespond - Vulnerability Risk Detection: leaderboard

Metric: Detection CAP-score (%) on the Vulnerability Risk (VUL) capability dimension: the detection (discovery, evidence and root-cause attribution, 0-3 per checkpoint) score achieved over all checkpoints mapped to the dimension across the 10 SecRespond cyber ranges (averaged over three LLM judges); agents in the OpenCode harness with reasoning enabled; higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 23 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Thinking)79.2
2Claude Sonnet 4.6 (Thinking)79.2
3GLM-5.178.5
4Claude Opus 4.7 (Thinking)76.2
5Claude Opus 4.5 (Thinking)72.6
6DeepSeek V3.2 (Thinking)69.6
7GLM-568.9
8Claude Sonnet 4.5 (Thinking)68.1
9Kimi K2.5 (Thinking)65.9
10MiniMax-M2.763.7
11Gemini 3 Flash62.2
12Qwen 3.7 Max (Thinking)62.2
13GPT-5.561.5
14GPT-5.4 Pro (xHigh)61.5
15Kimi K2.6 (Thinking)55.5

Interactive version: theaggregate.ai/benchmark?slug=secrespond-vulnerability-risk-detection · How It Works · Data refreshed daily, snapshot 2026-09-29.