CI-Repair-Bench - Fault Localization: leaderboard

Metric: Top-1 localization recall (%): share of the 567 real GitHub Actions CI failure instances from 103 Python repositories, repaired by the paper's fixed reference pipeline (CI log analysis, fault localization, patch generation) with only the LLM varied, and validated by full CI re-execution under the original workflow whose top-ranked predicted file is a ground-truth repair-relevant file, with the default agent-based log analysis; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 4 models tracked.

Top models

#ModelScore
1GPT-5 Mini45.33
2GPT-4o Mini (2024-07-18)42.5

Interactive version: theaggregate.ai/benchmark?slug=ci-repair-bench-fault-localization · How It Works · Data refreshed daily, snapshot 2026-10-07.