CITBench - Task Completion: leaderboard
Metric: Average task completion (%; share of repeated runs whose output table is fully correct, a Level-Accuracy of 1, on the offline single-turn tabular data processing tasks of CITBench, a balanced 50% slice of its 1,296 instances in 18 task types; temperature 1.0; unweighted mean of the Base, Complex-Rule, Multi-Input and Multi-Step tiers). Source: arxiv.org. Saturation forecast: Around February 2027. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 | 66.85 |
| 2 | Gemini 3 Flash | 60.12 |
| 3 | Gemini 3 Pro | 60.11 |
| 4 | GLM-4.7 | 57.62 |
| 5 | GPT-5.2 | 55.56 |
| 6 | Claude Haiku 4.5 | 50.43 |
| 7 | DeepSeek V3.2 | 49.17 |
| 8 | Qwen 3 235B A22B | 45.69 |
| 9 | Qwen 3 Coder Flash | 45.09 |
| 10 | GPT-4o | 43.65 |
| 11 | Qwen 3 32B | 42.97 |
| 12 | Qwen 3 14B | 37.15 |
Interactive version: theaggregate.ai/benchmark?slug=citbench-task-completion · How It Works · Data refreshed daily, snapshot 2026-09-26.