LLVM-Bench (Trae Agent): leaderboard
Metric: Resolved rate (%; share of the 423 validated LLVM issues, versions 18-21, whose generated patch applies, builds and passes the full LLVM test suite including the issue tests in LLVM-Gym; Trae Agent scaffold with at most 50 interaction turns, temperature 0, 8,192-token generation limit). Source: arxiv.org. Saturation forecast: Around 2030. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V3.2 | 8.98 |
| 2 | Grok Code Fast 1 | 7.33 |
| 3 | Gemini 3 Flash | 4.49 |
| 4 | Qwen 3 Coder Plus | 3.55 |
Interactive version: theaggregate.ai/benchmark?slug=llvm-bench-trae-agent · How It Works · Data refreshed daily, snapshot 2026-09-29.