CI-Repair-Bench (BM25 Log Retrieval): leaderboard
Metric: Pass@1 (%): share of the 567 real GitHub Actions CI failure instances from 103 Python repositories, repaired by the paper's fixed reference pipeline (CI log analysis, fault localization, patch generation) with only the LLM varied, and validated by full CI re-execution under the original workflow fixed by the first repair attempt when the log analysis stage uses BM25 retrieval of log segments instead of agent-based iterative analysis; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 Mini | 10.6 |
| 2 | GPT-4o Mini (2024-07-18) | 1.9 |
Interactive version: theaggregate.ai/benchmark?slug=ci-repair-bench-bm25-log-retrieval · How It Works · Data refreshed daily, snapshot 2026-10-07.