CLBench-V - Information Application (L1): leaderboard

Metric: Score (%; instance-weighted mean of the dataset scores, each 0-1 and scaled to percent, over the L1 new-information datasets (Pix2Fact, MMLongBench-Pic, BrowseComp-V3, Financial Report ROE, PRISMM-Bench); task-specific evaluators: option extraction, normalized or tolerance numeric match, a route-topology judge, and Qwen3.6-27B semantic or rubric judging for open-ended answers). Source: arxiv.org. Saturation forecast: Around July 2028. 6 models tracked.

Top models

#ModelScore
1Qwen 3.5 Plus29.54
2Qwen 3.6 27B28.78
3GPT-5.428.14
4Kimi K2.625.82
5Seed 2.0 Lite22.41

Interactive version: theaggregate.ai/benchmark?slug=clbench-v-information-application-l1 · How It Works · Data refreshed daily, snapshot 2026-09-29.