CLBench-V: leaderboard

Metric: Score (%; instance-weighted mean of the dataset scores, each 0-1 and scaled to percent, over all 3,443 instances of the 14 datasets; task-specific evaluators: option extraction, normalized or tolerance numeric match, a route-topology judge, and Qwen3.6-27B semantic or rubric judging for open-ended answers). Source: arxiv.org. Saturation forecast: Around 2028. 6 models tracked.

Top models

#ModelScore
1Kimi K2.619.91
2Qwen 3.6 27B19.58
3GPT-5.418.94
4Qwen 3.5 Plus18.62
5Seed 2.0 Lite18.5

Interactive version: theaggregate.ai/benchmark?slug=clbench-v · How It Works · Data refreshed daily, snapshot 2026-09-29.