CLBench — leaderboard

Knowledge-learning benchmark by Tencent evaluating LLM ability to learn from context across domain knowledge, rule systems, procedural tasks, and pattern discovery. 23 frontier models tested.

Metric: Solving Rate (%). Source: www.clbench.com. Status: saturation imminent. 36 models tracked.

Top models

#ModelScore
1GPT-5.4 (xHigh)27.9
2GPT-5.1 (High)23.7
3Hy3-preview22.8
4Grok 4.20 (Reasoning)22.2
5GPT-5.121.1
6Claude Opus 4.5 (Thinking)21.1
7Gemini 3.1 Pro (Preview) (High)20.8
8Claude Opus 4.620.7
9Qwen 3.6 Plus20.3
10Qwen 3.5 Plus (Thinking)19.8
11Kimi K2.519.3
12Claude Opus 4.519.1
13GLM-518.7
14GPT-5.218.2
15GPT-5.2 (High)18.1

Interactive version: theaggregate.ai/benchmark?slug=clbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.