CLBench-V - Grounding (L0): leaderboard

Metric: Score (%; instance-weighted mean of the dataset scores, each 0-1 and scaled to percent, over the L0 context-grounding datasets (ReasonMap, Insight-O3, ZeroBench, CourtSI, MIRBench); task-specific evaluators: option extraction, normalized or tolerance numeric match, a route-topology judge, and Qwen3.6-27B semantic or rubric judging for open-ended answers). Source: arxiv.org. Saturation forecast: Around 2028. 6 models tracked.

Top models

#ModelScore
1GPT-5.420.03
2Qwen 3.6 27B18.65
3Seed 2.0 Lite17.97
4Kimi K2.617.55
5Qwen 3.5 Plus15.07

Interactive version: theaggregate.ai/benchmark?slug=clbench-v-grounding-l0 · How It Works · Data refreshed daily, snapshot 2026-09-29.