LLVM-Bench (Live-SWE-agent): leaderboard

Metric: Resolved rate (%; share of the 423 validated LLVM issues, versions 18-21, whose generated patch applies, builds and passes the full LLVM test suite including the issue tests in LLVM-Gym; self-evolving Live-SWE-agent scaffold with at most 50 interaction turns, temperature 0, 8,192-token generation limit). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.

Top models

#ModelScore
1Live-SWE-agent + deepseek-v3.28.51
2Live-SWE-agent + grok-code-fast-17.57
3Live-SWE-agent + gemini-3-flash4.73
4Live-SWE-agent + qwen3-coder-plus2.6

Interactive version: theaggregate.ai/benchmark?slug=llvm-bench-live-swe-agent · How It Works · Data refreshed daily, snapshot 2026-09-29.