CI-Repair-Bench: leaderboard
Metric: Pass@1 (%): share of the 567 real GitHub Actions CI failure instances from 103 Python repositories, repaired by the paper's fixed reference pipeline (CI log analysis, fault localization, patch generation) with only the LLM varied, and validated by full CI re-execution under the original workflow that the first repair attempt fixes, with the default agent-based iterative log analysis; higher is better. Source: arxiv.org. Saturation forecast: Around May 2028. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 Mini | 18.9 |
| 2 | GPT-4o Mini (2024-07-18) | 7.9 |
Interactive version: theaggregate.ai/benchmark?slug=ci-repair-bench · How It Works · Data refreshed daily, snapshot 2026-10-07.