CCR-Bench - Workflow Task Success: leaderboard

Metric: Task success rate (0-1, times 100): share of tasks executed exactly as specified with every objective met, on CCR-Bench's 70 logical workflow control tasks (multi-turn tool use with conditional branching, implicit nested workflows, implicit tool calls and long tool chains, checked by verification scripts), temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro70#145
2GPT-4.152.9#240
3O3 Mini51.4#266
4DeepSeek R1 052840#217
5QwQ-32B38.6#410
6Qwen 3 32B (Thinking)38.6#424 (Qwen 3 32B)
7Qwen 2.5 72B Instruct32.9#436
8DeepSeek V3 (0324)30#332
9Qwen 3 32B (Non-reasoning)15.7#424 (Qwen 3 32B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=ccr-bench-workflow-task-success · How It Works · Data refreshed daily, snapshot 2026-10-11.