LLM2014 Logic 2026-06: leaderboard

Metric: Median Score. Source: raw.githubusercontent.com. 43 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)80.47
2Gemini 3.5 Flash (High)73.48
3GLM-5.2 (Max)71.16
4Claude Opus 4.8 (xHigh)68.32
5Gemini 3.1 Pro (Preview) (High)67.24
6DeepSeek V4 Pro (Max)63.93
7Seed 2.0 Pro (High)55.84
8MiniMax-M351.03
9DeepSeek V4 Flash (Max)49.41
10Kimi K2.647.58
11Hy3-preview46.12
12GPT-5.4 Mini (High)40.37
13Claude Opus 4.640.16
14Claude Sonnet 4.537.47
15Grok 4.331.63

Interactive version: theaggregate.ai/benchmark?slug=llm2014-logic-2026-06 · How It Works · Data refreshed daily, snapshot 2026-09-05.