CTI-REALM — leaderboard
Microsoft detection-engineering benchmark where agents interpret cyber threat intelligence, inspect telemetry, and produce Sigma/KQL detection rules across Linux, AKS, and Azure cloud scenarios.
Metric: Normalized Reward. Source: arxiv.org. Status: saturation imminent. 16 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 (High) | 0.64 |
| 2 | Claude Opus 4.5 | 0.62 |
| 3 | Claude Sonnet 4.5 | 0.59 |
| 4 | GPT-5 (Medium) | 0.57 |
| 5 | GPT-5.2 (Medium) | 0.57 |
| 6 | GPT-5 (High) | 0.56 |
| 7 | GPT-5.2 (High) | 0.55 |
| 8 | GPT-5 (Low) | 0.54 |
| 9 | GPT-5.1 (Medium) | 0.51 |
| 10 | GPT-5.1 (High) | 0.49 |
| 11 | GPT-5.2 (Low) | 0.49 |
| 12 | O3 | 0.47 |
| 13 | GPT-5 Mini | 0.45 |
| 14 | GPT-4.1 | 0.42 |
| 15 | GPT-5.1 (Low) | 0.37 |
Interactive version: theaggregate.ai/benchmark?slug=cti-realm · How the rankings work · Data refreshed daily, snapshot 2026-07-22.