SecRespond - Baseline Risk Planning: leaderboard

Metric: Planning CAP-score (%) on the Baseline Risk (BAS) capability dimension: the remediation-planning (correct and complete fix, 0-2 per checkpoint) score achieved over all checkpoints mapped to the dimension across the 10 SecRespond cyber ranges (averaged over three LLM judges); agents in the OpenCode harness with reasoning enabled; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 23 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.5 (Thinking)70.8
2Claude Sonnet 4.6 (Thinking)69.8
3Claude Opus 4.6 (Thinking)67.6
4GLM-5.165.7
5Qwen 3.7 Max (Thinking)63.7
6Claude Opus 4.5 (Thinking)63.5
7Qwen 3.6 Plus (Thinking)62.6
8GLM-554.1
9Kimi K2.5 (Thinking)52.9
10GPT-5.4 Pro (xHigh)51.9
11MiniMax-M2.751.2
12DeepSeek V3.2 (Thinking)50
13Qwen 3.5 Plus (Thinking)49
14Kimi K2.6 (Thinking)48.6
15GPT-5.546.8

Interactive version: theaggregate.ai/benchmark?slug=secrespond-baseline-risk-planning · How It Works · Data refreshed daily, snapshot 2026-09-29.