CI-Repair-Bench (BM25 Log Retrieval): leaderboard

Metric: Pass@1 (%): share of the 567 real GitHub Actions CI failure instances from 103 Python repositories, repaired by the paper's fixed reference pipeline (CI log analysis, fault localization, patch generation) with only the LLM varied, and validated by full CI re-execution under the original workflow fixed by the first repair attempt when the log analysis stage uses BM25 retrieval of log segments instead of agent-based iterative analysis; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 4 models tracked.

Top models

#ModelScore
1GPT-5 Mini10.6
2GPT-4o Mini (2024-07-18)1.9

Interactive version: theaggregate.ai/benchmark?slug=ci-repair-bench-bm25-log-retrieval · How It Works · Data refreshed daily, snapshot 2026-10-07.