LLM2014 Logic 2026-08: leaderboard

Metric: Median Score. Source: raw.githubusercontent.com. 46 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)73.89
2GPT-5.6 Sol (xHigh)69.42
3Kimi K3 (Max)67.66
4Claude Opus 5 (xHigh)64.74
5GLM-5.3 (Max)63.91
6GLM-5.3 Flash (Max)60.52
7Qwen 3.8 Max (xHigh)60.05
8DeepSeek V4 Pro (0813) (Max)59.63
9Gemini 3.7 Flash (High)57.56
10Grok 4.6 (High)56.41
11Gemini 3.1 Pro (Preview) (High)55.84
12DeepSeek V4 Flash (0731) (Max)55.23
13GLM-5.2 (Max)54.7
14GPT-5.6 Luna (xHigh)51.62
15Qwen 3.8 27B (xHigh)47.65

Interactive version: theaggregate.ai/benchmark?slug=llm2014-logic-2026-08 · How It Works · Data refreshed daily, snapshot 2026-09-05.