SecRespond - Persistence Mechanism Planning: leaderboard

Metric: Planning CAP-score (%) on the Persistence Mechanism (PER) capability dimension: the remediation-planning (correct and complete fix, 0-2 per checkpoint) score achieved over all checkpoints mapped to the dimension across the 10 SecRespond cyber ranges (averaged over three LLM judges); agents in the OpenCode harness with reasoning enabled; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 23 models tracked.

Top models

#ModelScore
1GLM-5.146.7
2Qwen 3.7 Max (Thinking)46.7
3Claude Opus 4.6 (Thinking)45.8
4Claude Sonnet 4.6 (Thinking)44.9
5Claude Sonnet 4.5 (Thinking)44.1
6Claude Opus 4.5 (Thinking)44
7GPT-5.2 Pro39
8GLM-536.6
9Kimi K2.5 (Thinking)36
10GPT-5.4 Pro (xHigh)35.2
11Kimi K2.6 (Thinking)33.9
12Qwen 3.6 Plus (Thinking)32.2
13Qwen 3.5 Plus (Thinking)31.5
14MiniMax-M2.730.3
15DeepSeek V3.2 (Thinking)30.1

Interactive version: theaggregate.ai/benchmark?slug=secrespond-persistence-mechanism-planning · How It Works · Data refreshed daily, snapshot 2026-09-29.