LLM2014 Logic 2026-06 — leaderboard

Metric: Median Score. Source: raw.githubusercontent.com. 43 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)80.47
2Gemini 3.5 Flash (High)73.48
3GLM-5.2 (Max)71.16
4Claude Opus 4.8 (xHigh)68.32
5Gemini 3.1 Pro (Preview) (High)67.24
6DeepSeek V4 Pro (Max)63.93
7MiniMax-M351.03
8DeepSeek V4 Flash (Max)49.41
9Kimi K2.647.58
10Hy3-preview46.12
11GPT-5.4 Mini (High)40.37
12Claude Opus 4.640.16
13Claude Sonnet 4.537.47
14Grok 4.331.63
15Gemini 3.5 Flash (Minimal)31.45

Interactive version: theaggregate.ai/benchmark?slug=llm2014-logic-2026-06 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.