CAPA - First-Turn Success: leaderboard
Metric: First-turn executable success (%; share of the 300 held-out coding sessions whose first assistant turn is an accepted submission, with no clarification asked; same-user history: the model also sees the dialogue traces of five resolved sessions from the same user, whose recurring ambiguity pattern the held-out request repeats). Source: arxiv.org. Saturation forecast: Around 2030. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.8 | 60.3 |
| 2 | GLM-5.2 | 46.7 |
| 3 | Gemini 3.5 Flash | 38.7 |
| 4 | Kimi K2.6 | 32.3 |
| 5 | GPT-5.5 | 31 |
| 6 | Qwen 3.7 Max | 29.7 |
| 7 | Qwen 3 8B | 23 |
| 8 | Qwen 3.5 27B | 21.3 |
| 9 | GPT-5.6 Sol | 18.3 |
| 10 | DeepSeek V4 Pro | 17.7 |
| 11 | Llama 3.3 70B Instruct | 16 |
| 12 | Claude Sonnet 4.6 | 14 |
Interactive version: theaggregate.ai/benchmark?slug=capa-first-turn-success · How It Works · Data refreshed daily, snapshot 2026-09-29.