CAPA - First-Turn Success: leaderboard

Metric: First-turn executable success (%; share of the 300 held-out coding sessions whose first assistant turn is an accepted submission, with no clarification asked; same-user history: the model also sees the dialogue traces of five resolved sessions from the same user, whose recurring ambiguity pattern the held-out request repeats). Source: arxiv.org. Saturation forecast: Around 2030. 12 models tracked.

Top models

#ModelScore
1Claude Opus 4.860.3
2GLM-5.246.7
3Gemini 3.5 Flash38.7
4Kimi K2.632.3
5GPT-5.531
6Qwen 3.7 Max29.7
7Qwen 3 8B23
8Qwen 3.5 27B21.3
9GPT-5.6 Sol18.3
10DeepSeek V4 Pro17.7
11Llama 3.3 70B Instruct16
12Claude Sonnet 4.614

Interactive version: theaggregate.ai/benchmark?slug=capa-first-turn-success · How It Works · Data refreshed daily, snapshot 2026-09-29.