CLBench: leaderboard

Knowledge-learning benchmark by Tencent evaluating LLM ability to learn from context across domain knowledge, rule systems, procedural tasks, and pattern discovery. 23 frontier models tested.

Metric: Solving Rate (%). Source: www.clbench.com. Status: saturation imminent. 36 models tracked.

Top models

#ModelScore
1GPT-5.4 (xHigh)27.9
2GPT-5.1 (High)23.7
3Hy3-preview22.8
4Grok 4.20 (Reasoning)22.2
5GPT-5.121.1
6Claude Opus 4.5 (Thinking)21.1
7Gemini 3.1 Pro (Preview) (High)20.8
8Claude Opus 4.620.7
9Seed 2.0 Pro (High)20.6
10Qwen 3.6 Plus20.3
11Qwen 3.5 Plus (Thinking)19.8
12Kimi K2.519.3
13Claude Opus 4.519.1
14GLM-518.7
15GPT-5.218.2

Interactive version: theaggregate.ai/benchmark?slug=clbench · How It Works · Data refreshed daily, snapshot 2026-09-05.