SecRespond - Vulnerability Risk Planning: leaderboard

Metric: Planning CAP-score (%) on the Vulnerability Risk (VUL) capability dimension: the remediation-planning (correct and complete fix, 0-2 per checkpoint) score achieved over all checkpoints mapped to the dimension across the 10 SecRespond cyber ranges (averaged over three LLM judges); agents in the OpenCode harness with reasoning enabled; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 23 models tracked.

Top models

#ModelScore
1Claude Opus 4.7 (Thinking)72.6
2GLM-5.171.1
3Claude Sonnet 4.6 (Thinking)67.8
4Claude Sonnet 4.5 (Thinking)66.7
5Claude Opus 4.6 (Thinking)63.3
6Kimi K2.5 (Thinking)60
7GLM-558.8
8DeepSeek V3.2 (Thinking)54.5
9Claude Opus 4.5 (Thinking)54.4
10Qwen 3.6 Plus (Thinking)52.3
11Qwen 3.7 Max (Thinking)47.8
12GPT-5.4 Pro (xHigh)45.6
13Qwen 3.5 Plus (Thinking)44.4
14Gemini 3 Flash43.3
15Kimi K2.6 (Thinking)41.1

Interactive version: theaggregate.ai/benchmark?slug=secrespond-vulnerability-risk-planning · How It Works · Data refreshed daily, snapshot 2026-09-29.