CAPA (No History): leaderboard

Metric: Executable success (%; share of the 300 held-out coding sessions solved within an eight-turn budget, a submission counting when the execution-grounded judge accepts it from the hidden tests and the reference solution, and a simulated user agent answering clarification questions; no history: the model sees only the new ambiguous request). Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.

Top models

#ModelScore
1Claude Opus 4.888
2GLM-5.285.3
3Gemini 3.5 Flash79.3
4GPT-5.6 Sol79
5Qwen 3.7 Max77.7
6Kimi K2.677
7DeepSeek V4 Pro76.3
8Claude Sonnet 4.676
9GPT-5.574.3
10Qwen 3.5 27B56
11Qwen 3 8B23.3
12Llama 3.3 70B Instruct10

Interactive version: theaggregate.ai/benchmark?slug=capa-no-history · How It Works · Data refreshed daily, snapshot 2026-09-29.