CCTU (Non-Thinking) - Solve Rate: leaderboard
Metric: Solve rate (%): the share of cases in which the model solves every subquery with all constraints satisfied at the end (violations repaired after validator feedback count), on the 200 CCTU test cases (FTRL tool-use tasks rewritten with on average seven constraints from 12 categories over resources, behavior, toolsets and responses; locally executable tools), an executable validator checking every step and returning violation feedback for revision, thinking mode disabled, default API settings otherwise, mean of 3 runs; higher is better. Source: arxiv.org. Saturation forecast: Around October 2028. 9 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Claude Opus 4.6 | 34.5 | #60 |
| 2 | Kimi K2.5 (Non-reasoning) | 22.67 | #139 (Kimi K2.5) |
| 3 | Qwen 3.5 Plus (Non-reasoning) | 21.33 | #123 (Qwen 3.5 Plus) |
| 4 | GPT-5.2 (Non-reasoning) | 20.33 | #105 (GPT-5.2) |
| 5 | Gemini 3 Pro | 19 | #77 |
| 6 | GPT-5.1 (Non-reasoning) | 18.17 | #131 (GPT-5.1) |
| 7 | Seed 2.0 Pro (Non-reasoning) | 18.17 | #111 (Seed 2.0 Pro) |
| 8 | DeepSeek V3.2 (Non-reasoning) | 17 | #198 (DeepSeek V3.2) |
| 9 | O3 | 11.5 | #121 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=cctu-non-thinking-solve-rate · How It Works · Data refreshed daily, snapshot 2026-10-11.