CCR-Bench - Workflow Task Completion: leaderboard

Metric: Task completion rate (0-1, times 100): share of required tool invocations or path steps completed, averaged over tasks, on CCR-Bench's 70 logical workflow control tasks (multi-turn tool use with conditional branching, implicit nested workflows, implicit tool calls and long tool chains, checked by verification scripts), temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro84.4#145
2GPT-4.180#240
3O3 Mini76.8#266
4QwQ-32B69.3#410
5Qwen 3 32B (Thinking)65.7#424 (Qwen 3 32B)
6DeepSeek R1 052864.4#217
7Qwen 2.5 72B Instruct63.1#436
8DeepSeek V3 (0324)56.2#332
9Qwen 3 32B (Non-reasoning)43.8#424 (Qwen 3 32B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=ccr-bench-workflow-task-completion · How It Works · Data refreshed daily, snapshot 2026-10-11.