LLVM-Bench (Live-SWE-agent): leaderboard
Metric: Resolved rate (%; share of the 423 validated LLVM issues, versions 18-21, whose generated patch applies, builds and passes the full LLVM test suite including the issue tests in LLVM-Gym; self-evolving Live-SWE-agent scaffold with at most 50 interaction turns, temperature 0, 8,192-token generation limit). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Live-SWE-agent + deepseek-v3.2 | 8.51 |
| 2 | Live-SWE-agent + grok-code-fast-1 | 7.57 |
| 3 | Live-SWE-agent + gemini-3-flash | 4.73 |
| 4 | Live-SWE-agent + qwen3-coder-plus | 2.6 |
Interactive version: theaggregate.ai/benchmark?slug=llvm-bench-live-swe-agent · How It Works · Data refreshed daily, snapshot 2026-09-29.