CAPA: leaderboard

Metric: Executable success (%; share of the 300 held-out coding sessions solved within an eight-turn budget, a submission counting when the execution-grounded judge accepts it from the hidden tests and the reference solution, and a simulated user agent answering clarification questions; same-user history: the model also sees the dialogue traces of five resolved sessions from the same user, whose recurring ambiguity pattern the held-out request repeats). Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.

Top models

#ModelScore
1Claude Opus 4.890
2GLM-5.289.7
3GPT-5.584.3
4Claude Sonnet 4.683.7
5Qwen 3.7 Max82.7
6Kimi K2.682
7Gemini 3.5 Flash81.7
8DeepSeek V4 Pro79
9GPT-5.6 Sol78.7
10Qwen 3.5 27B74.3
11Llama 3.3 70B Instruct30
12Qwen 3 8B27.7

Interactive version: theaggregate.ai/benchmark?slug=capa · How It Works · Data refreshed daily, snapshot 2026-09-29.