LLVM-Bench (SWE-agent): leaderboard

Metric: Resolved rate (%; share of the 423 validated LLVM issues, versions 18-21, whose generated patch applies, builds and passes the full LLVM test suite including the issue tests in LLVM-Gym; SWE-agent scaffold with at most 50 interaction turns, temperature 0, 8,192-token generation limit). Source: arxiv.org. Saturation forecast: Around 2031. 4 models tracked.

Top models

#ModelScore
1Grok Code Fast 110.87
2DeepSeek V3.27.8
3Gemini 3 Flash4.26
4Qwen 3 Coder Plus3.78

Interactive version: theaggregate.ai/benchmark?slug=llvm-bench-swe-agent · How It Works · Data refreshed daily, snapshot 2026-09-29.