SecRespond - Planning: leaderboard

Metric: Planning CHK-score (%): the remediation-planning (correct and complete fix, 0-2 per checkpoint) score achieved over each range checkpoints (averaged over Claude Opus 4.7, Gemini 3.1 Pro and GPT-5.4 Pro as judges), averaged over the 10 SecRespond cyber ranges; agents in the OpenCode harness with reasoning enabled investigate a forensic disk snapshot of a compromised host and write forensic reports and a remediation plan; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 23 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.5 (Thinking)61.2
2GLM-5.159.2
3Claude Sonnet 4.6 (Thinking)59.1
4Claude Opus 4.6 (Thinking)58
5Claude Opus 4.5 (Thinking)56.4
6Qwen 3.7 Max (Thinking)54.3
7Kimi K2.5 (Thinking)50.7
8Qwen 3.6 Plus (Thinking)48.3
9GLM-545.9
10Qwen 3.5 Plus (Thinking)44
11DeepSeek V3.2 (Thinking)43.6
12Kimi K2.6 (Thinking)43
13GPT-5.4 Pro (xHigh)40.7
14MiniMax-M2.740
15GPT-5.2 Pro37.1

Interactive version: theaggregate.ai/benchmark?slug=secrespond-planning · How It Works · Data refreshed daily, snapshot 2026-09-29.