SecRespond - Intrusion Entity Planning: leaderboard

Metric: Planning CAP-score (%) on the Intrusion Entity (ENT) capability dimension: the remediation-planning (correct and complete fix, 0-2 per checkpoint) score achieved over all checkpoints mapped to the dimension across the 10 SecRespond cyber ranges (averaged over three LLM judges); agents in the OpenCode harness with reasoning enabled; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 23 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.6 (Thinking)62.7
2GLM-5.161.8
3Claude Opus 4.5 (Thinking)60.8
4Claude Sonnet 4.5 (Thinking)58.4
5Claude Opus 4.6 (Thinking)58.1
6Kimi K2.5 (Thinking)52.4
7Qwen 3.7 Max (Thinking)52.4
8Qwen 3.6 Plus (Thinking)47.1
9Qwen 3.5 Plus (Thinking)44.6
10Kimi K2.6 (Thinking)43.9
11DeepSeek V3.2 (Thinking)42.9
12GLM-542.1
13GPT-5.4 Pro (xHigh)38.3
14GPT-5.2 Pro36
15Gemini 3 Flash35.3

Interactive version: theaggregate.ai/benchmark?slug=secrespond-intrusion-entity-planning · How It Works · Data refreshed daily, snapshot 2026-09-29.