From 0-Order Selection to 2-Order Judgment — leaderboard

Metric: H-Comb (self-reported). Source: benchmarklist.com. 13 models tracked.

Top models

#ModelScore
1GLM-538.33
2GPT-5.431.67
3DeepSeek V4 Pro30
4GLM-4.730
5Claude Opus 4.626.67
6O325
7Gemini 3.1 Pro (Preview)23.33
8Qwen 3.5 Plus (2026-04-20)18.33
9DeepSeek V3.216.67
10Qwen 3.6 Plus5

Interactive version: theaggregate.ai/benchmark?slug=from-0-order-selection-to-2-order-judgment · How the rankings work · Data refreshed daily, snapshot 2026-07-22.