CCTU (Thinking) - Solve Rate: leaderboard

Metric: Solve rate (%): the share of cases in which the model solves every subquery with all constraints satisfied at the end (violations repaired after validator feedback count), on the 200 CCTU test cases (FTRL tool-use tasks rewritten with on average seven constraints from 12 categories over resources, behavior, toolsets and responses; locally executable tools), an executable validator checking every step and returning violation feedback for revision, thinking mode enabled, default API settings otherwise, mean of 3 runs; higher is better. Source: arxiv.org. Saturation forecast: Around January 2028. 9 models tracked.

Top models

#ModelScoreOverall rank
1Claude Opus 4.6 (Thinking)34.17#60 (Claude Opus 4.6)
2Qwen 3.5 Plus (Thinking)24.83#123 (Qwen 3.5 Plus)
3GPT-5.2 (Thinking)24.5#105 (GPT-5.2)
4GPT-5.1 (Thinking)22.83#131 (GPT-5.1)
5Kimi K2.5 (Thinking)21.33#139 (Kimi K2.5)
6Seed 2.0 Pro (Thinking)20.33#111 (Seed 2.0 Pro)
7Gemini 3 Pro19.33#77
8DeepSeek V3.2 (Thinking)18#198 (DeepSeek V3.2)
9O311.83#121

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=cctu-thinking-solve-rate · How It Works · Data refreshed daily, snapshot 2026-10-11.