LLM2014 Logic 2026-05 — leaderboard

Metric: Median Score. Source: raw.githubusercontent.com. 43 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)80.47
2Claude Opus 4.6 (High)76.48
3Gemini 3.5 Flash (High)73.93
4Claude Opus 4.8 (xHigh)68.32
5Gemini 3.1 Pro (Preview) (High)67.91
6DeepSeek V4 Pro (Max)64.82
7GLM-5.155.59
8DeepSeek V4 Flash (Max)50.97
9Kimi K2.647.81
10Hy3-preview47.01
11Qwen 3.6 Plus44.99
12Claude Opus 4.643.29
13GPT-5.4 Mini (High)40.59
14Claude Sonnet 4.539.93
15MiniMax-M2.735.94

Interactive version: theaggregate.ai/benchmark?slug=llm2014-logic-2026-05 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.