SecRespond - Investigation and Response Quality Detection: leaderboard

Metric: Detection CAP-score (%) on the Investigation and Response Quality (Q) capability dimension: the detection (discovery, evidence and root-cause attribution, 0-3 per checkpoint) score achieved over all checkpoints mapped to the dimension across the 10 SecRespond cyber ranges (averaged over three LLM judges); agents in the OpenCode harness with reasoning enabled; higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 23 models tracked.

Top models

#ModelScore
1GLM-5.175.5
2Claude Sonnet 4.5 (Thinking)71.6
3Qwen 3.7 Max (Thinking)69.1
4Claude Opus 4.6 (Thinking)68.8
5Claude Sonnet 4.6 (Thinking)66.7
6GPT-5.4 Pro (xHigh)64
7GPT-5.2 Pro62.9
8Claude Opus 4.5 (Thinking)62.4
9GPT-5.561.4
10Qwen 3.6 Plus (Thinking)58
11GLM-556.5
12Kimi K2.6 (Thinking)55.7
13GPT-5.455.2
14Kimi K2.5 (Thinking)54.9
15DeepSeek V3.2 (Thinking)54.8

Interactive version: theaggregate.ai/benchmark?slug=secrespond-investigation-and-response-quality-detection · How It Works · Data refreshed daily, snapshot 2026-09-29.