CTI-REALM — leaderboard

Microsoft detection-engineering benchmark where agents interpret cyber threat intelligence, inspect telemetry, and produce Sigma/KQL detection rules across Linux, AKS, and Azure cloud scenarios.

Metric: Normalized Reward. Source: arxiv.org. Status: saturation imminent. 16 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (High)0.64
2Claude Opus 4.50.62
3Claude Sonnet 4.50.59
4GPT-5 (Medium)0.57
5GPT-5.2 (Medium)0.57
6GPT-5 (High)0.56
7GPT-5.2 (High)0.55
8GPT-5 (Low)0.54
9GPT-5.1 (Medium)0.51
10GPT-5.1 (High)0.49
11GPT-5.2 (Low)0.49
12O30.47
13GPT-5 Mini0.45
14GPT-4.10.42
15GPT-5.1 (Low)0.37

Interactive version: theaggregate.ai/benchmark?slug=cti-realm · How the rankings work · Data refreshed daily, snapshot 2026-07-22.