SecIT Bench (Pydantic AI): leaderboard

Cribl's benchmark for IT and security operations agents, run under a plain Pydantic AI harness. Thirty telemetry scenarios across intrusion and breach, covert egress, service errors and performance degradation are each rolled out three times; an incomplete answer scores zero, and the reported figure is the mean scenario accuracy.

Metric: Accuracy (%). Source: secitbench.cribl.io. Status: saturation imminent. 18 models tracked.

Top models

#ModelScore
1GLM-5.381.44
2Grok 4.680.65
3Claude Opus 579.15
4Qwen 3.8 Max78.52
5Muse Spark 1.377.53
6Kimi K376.99
7GLM-5.3 Flash75.48
8GPT-5.6 Sol75.4
9Grok 4.574.88
10GPT-5.6 Terra69.38
11Muse Spark 1.269.35
12DeepSeek V4 Flash (0731)69.02
13Gemini 3.8 Flash69
14Claude Opus 4.867.04
15Gemini 3.7 Flash65.12

Interactive version: theaggregate.ai/benchmark?slug=secit-bench-pydantic-ai · How It Works · Data refreshed daily, snapshot 2026-09-05.