ComBench (Compilation Error Repair, OpenHands Agent): leaderboard

Metric: Semantic correctness (%) of patches for the 200 reproducible real CI compilation errors from Bitcoin, OpenSSL, RocksDB and LLVM: the patch compiles without new errors and, by the authors' manual analysis, matches the developer's intended semantics, agentic repair in the OpenHands framework (version 0.55.0, at most 100 iterations): the model starts from the raw compiler error log only and must localize and fix the error by inspecting files and running builds (one retry when no legal diff is produced), temperature 0.2; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 2 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 442#194
2Qwen 3 32B19#424

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=combench-compilation-error-repair-openhands-agent · How It Works · Data refreshed daily, snapshot 2026-10-11.