CITBench - Task Completion: leaderboard

Metric: Average task completion (%; share of repeated runs whose output table is fully correct, a Level-Accuracy of 1, on the offline single-turn tabular data processing tasks of CITBench, a balanced 50% slice of its 1,296 instances in 18 task types; temperature 1.0; unweighted mean of the Base, Complex-Rule, Multi-Input and Multi-Step tiers). Source: arxiv.org. Saturation forecast: Around February 2027. 12 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.566.85
2Gemini 3 Flash60.12
3Gemini 3 Pro60.11
4GLM-4.757.62
5GPT-5.255.56
6Claude Haiku 4.550.43
7DeepSeek V3.249.17
8Qwen 3 235B A22B45.69
9Qwen 3 Coder Flash45.09
10GPT-4o43.65
11Qwen 3 32B42.97
12Qwen 3 14B37.15

Interactive version: theaggregate.ai/benchmark?slug=citbench-task-completion · How It Works · Data refreshed daily, snapshot 2026-09-26.