OpenRCA — leaderboard

Root cause analysis benchmark from Microsoft. Tests LLMs on diagnosing software failures from log and telemetry data across 335 real failure cases with 68 GB of telemetry.

Metric: Accuracy (%). Source: microsoft.github.io. Status: saturation imminent. 10 models tracked.

Top models

#ModelScore
1Claude Opus 4.636.42
2Claude Opus 4.528.36
3GPT-5.219.4
4Gemini 3 Pro12.54
5Claude 3.5 Sonnet11.34
6GPT-4o8.96
7Gemini 1.5 Pro7.16
8Command-R+4.78
9Mistral Large 2 (Jul)4.48

Interactive version: theaggregate.ai/benchmark?slug=openrca · How the rankings work · Data refreshed daily, snapshot 2026-07-22.