SecRespond - Detection: leaderboard

Metric: Detection CHK-score (%): the detection (discovery, evidence and root-cause attribution, 0-3 per checkpoint) score achieved over each range checkpoints (averaged over Claude Opus 4.7, Gemini 3.1 Pro and GPT-5.4 Pro as judges), averaged over the 10 SecRespond cyber ranges; agents in the OpenCode harness with reasoning enabled investigate a forensic disk snapshot of a compromised host and write forensic reports and a remediation plan; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 23 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Thinking)78.2
2Qwen 3.7 Max (Thinking)78
3GLM-5.176.3
4Claude Sonnet 4.6 (Thinking)74.7
5GPT-5.4 Pro (xHigh)72.5
6GPT-5.570.7
7Claude Sonnet 4.5 (Thinking)69.7
8Claude Opus 4.5 (Thinking)69.5
9Qwen 3.6 Plus (Thinking)68.4
10GPT-5.2 Pro66.2
11Kimi K2.6 (Thinking)65.8
12GLM-564.9
13Qwen 3.5 Plus (Thinking)63.5
14Kimi K2.5 (Thinking)62.1
15GPT-5.460.7

Interactive version: theaggregate.ai/benchmark?slug=secrespond-detection · How It Works · Data refreshed daily, snapshot 2026-09-29.