SecRespond - Investigation and Response Quality Planning: leaderboard

Metric: Planning CAP-score (%) on the Investigation and Response Quality (Q) capability dimension: the remediation-planning (correct and complete fix, 0-2 per checkpoint) score achieved over all checkpoints mapped to the dimension across the 10 SecRespond cyber ranges (averaged over three LLM judges); agents in the OpenCode harness with reasoning enabled; higher is better. Source: arxiv.org. Saturation forecast: Around 2035. 23 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.5 (Thinking)66.7
2Kimi K2.5 (Thinking)42.3
3GLM-5.141.1
4DeepSeek V3.2 (Thinking)36.7
5GPT-5.4 Pro (xHigh)35.5
6Claude Opus 4.6 (Thinking)34.4
7Claude Sonnet 4.6 (Thinking)34.4
8Claude Opus 4.5 (Thinking)33.3
9MiniMax-M2.732.3
10Qwen 3.7 Max (Thinking)32.2
11Qwen 3.5 Plus (Thinking)30
12GPT-5.2 Pro27.8
13Kimi K2.6 (Thinking)26.7
14GLM-525.5
15Qwen 3.6 Plus (Thinking)24.5

Interactive version: theaggregate.ai/benchmark?slug=secrespond-investigation-and-response-quality-planning · How It Works · Data refreshed daily, snapshot 2026-09-29.