CLBench - Rule System Application: leaderboard

Metric: Solving Rate (%). Source: www.clbench.com. 36 models tracked.

Top models

#ModelScore
1GPT-5.4 (xHigh)26.7
2GPT-5.1 (High)23.7
3Grok 4.20 (Reasoning)22.7
4Hy3-preview21.6
5Claude Opus 4.621.4
6GPT-5.121
7Seed 2.0 Pro (High)20.7
8Gemini 3.1 Pro (Preview) (High)20.2
9Qwen 3.6 Plus19.6
10Seed 2.0 Pro (Medium)19.3
11Claude Opus 4.519.1
12Claude Opus 4.5 (Thinking)19
13Qwen 3.5 Plus (Thinking)18.9
14GPT-5.218
15Kimi K2.517.9

Interactive version: theaggregate.ai/benchmark?slug=clbench-rule-system-application · How It Works · Data refreshed daily, snapshot 2026-09-19.