CLBench-V - Knowledge Learning (L2): leaderboard

Metric: Score (%; instance-weighted mean of the dataset scores, each 0-1 and scaled to percent, over the L2 new-knowledge datasets (CL-Bench-Table, Paper Conclusion); task-specific evaluators: option extraction, normalized or tolerance numeric match, a route-topology judge, and Qwen3.6-27B semantic or rubric judging for open-ended answers). Source: arxiv.org. Saturation forecast: Around 2029. 6 models tracked.

Top models

#ModelScore
1Kimi K2.621.86
2Qwen 3.5 Plus19.64
3Seed 2.0 Lite16.55
4Qwen 3.6 27B13.98
5GPT-5.46.94

Interactive version: theaggregate.ai/benchmark?slug=clbench-v-knowledge-learning-l2 · How It Works · Data refreshed daily, snapshot 2026-09-29.