ITBench-AA: leaderboard

Artificial Analysis implementation of IBM's ITBench SRE benchmark, testing AI agents on Kubernetes incident root-cause analysis from offline incident snapshots.

Metric: Average Precision at Full Recall (self-reported). Source: benchmarklist.com. Status: saturation imminent. 23 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct60
2Claude Opus 4.7 (Adaptive Reasoning, Max Effort)46.7
3GPT-5.5 (xHigh)45.8
4Qwen 3.7 Max42.5
5Gemini 3.5 Flash (High)40.3
6GLM-5.1 (Thinking)40.3
7Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)39.8
8DeepSeek V4 Pro (Reasoning, Max Effort)38.3
9MiMo-V2.5-Pro38.2
10Gemma 4 31B (Reasoning)37.3
11GPT-5.4 Mini (xHigh)35.2
12GPT-5.4 (xHigh)34.5
13Qwen 3.5 397B A17B (Reasoning)34.1
14Grok 4.3 (High)32.7
15DeepSeek V4 Flash (Reasoning, Max Effort)31.5

Interactive version: theaggregate.ai/benchmark?slug=itbench-aa · How It Works · Data refreshed daily, snapshot 2026-09-05.