SecRespond - Intrusion Entity Detection: leaderboard

Metric: Detection CAP-score (%) on the Intrusion Entity (ENT) capability dimension: the detection (discovery, evidence and root-cause attribution, 0-3 per checkpoint) score achieved over all checkpoints mapped to the dimension across the 10 SecRespond cyber ranges (averaged over three LLM judges); agents in the OpenCode harness with reasoning enabled; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 23 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Thinking)86
2GLM-5.184.1
3Qwen 3.7 Max (Thinking)84.1
4Claude Sonnet 4.6 (Thinking)81.6
5Qwen 3.5 Plus (Thinking)80.2
6GPT-5.578.4
7GPT-5.4 Pro (xHigh)78.3
8Qwen 3.6 Plus (Thinking)77.1
9Claude Opus 4.5 (Thinking)76.5
10GPT-5.2 Pro73.8
11Claude Sonnet 4.5 (Thinking)72.8
12Kimi K2.6 (Thinking)72.6
13Kimi K2.5 (Thinking)72
14GPT-5.469.9
15DeepSeek V3.2 (Thinking)69.1

Interactive version: theaggregate.ai/benchmark?slug=secrespond-intrusion-entity-detection · How It Works · Data refreshed daily, snapshot 2026-09-29.