CLBench-V - Information Application (L1): leaderboard
Metric: Score (%; instance-weighted mean of the dataset scores, each 0-1 and scaled to percent, over the L1 new-information datasets (Pix2Fact, MMLongBench-Pic, BrowseComp-V3, Financial Report ROE, PRISMM-Bench); task-specific evaluators: option extraction, normalized or tolerance numeric match, a route-topology judge, and Qwen3.6-27B semantic or rubric judging for open-ended answers). Source: arxiv.org. Saturation forecast: Around July 2028. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3.5 Plus | 29.54 |
| 2 | Qwen 3.6 27B | 28.78 |
| 3 | GPT-5.4 | 28.14 |
| 4 | Kimi K2.6 | 25.82 |
| 5 | Seed 2.0 Lite | 22.41 |
Interactive version: theaggregate.ai/benchmark?slug=clbench-v-information-application-l1 · How It Works · Data refreshed daily, snapshot 2026-09-29.