LLM2014 Logic 2026-07 — leaderboard

Metric: Median Score. Source: raw.githubusercontent.com. 42 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)77.46
2Kimi K3 (Max)74.8
3GPT-5.6 Sol (xHigh)72.99
4Claude Opus 4.8 (xHigh)66.7
5Gemini 3.5 Flash (High)60.39
6Gemini 3.1 Pro (Preview) (High)58.46
7GLM-5.2 (Max)58.27
8Muse Spark 1.157.23
9Grok 4.5 (High)56.72
10GPT-5.6 Luna (xHigh)51.97
11DeepSeek V4 Pro (Max)49.97
12Claude Sonnet 5 (xHigh)43.14
13MiniMax-M340.94
14DeepSeek V4 Flash (Max)36.24
15Claude Opus 4.625.79

Interactive version: theaggregate.ai/benchmark?slug=llm2014-logic-2026-07 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.