SecRespond - Baseline Risk Detection: leaderboard

Metric: Detection CAP-score (%) on the Baseline Risk (BAS) capability dimension: the detection (discovery, evidence and root-cause attribution, 0-3 per checkpoint) score achieved over all checkpoints mapped to the dimension across the 10 SecRespond cyber ranges (averaged over three LLM judges); agents in the OpenCode harness with reasoning enabled; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 23 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.6 (Thinking)84.2
2GLM-5.181.1
3Qwen 3.7 Max (Thinking)80.9
4Claude Opus 4.6 (Thinking)78.6
5Claude Opus 4.5 (Thinking)75.8
6Qwen 3.6 Plus (Thinking)74.7
7GPT-5.4 Pro (xHigh)73.5
8MiniMax-M2.768.6
9GLM-567.8
10Kimi K2.6 (Thinking)66.7
11Claude Sonnet 4.5 (Thinking)66.2
12GPT-5.563.6
13GPT-5.2 Pro62.4
14Qwen 3.5 Plus (Thinking)61.9
15GPT-5.461.8

Interactive version: theaggregate.ai/benchmark?slug=secrespond-baseline-risk-detection · How It Works · Data refreshed daily, snapshot 2026-09-29.