LLM2014 Logic 2025-04 — leaderboard

Metric: Median Score. Source: raw.githubusercontent.com. 29 models tracked.

Top models

#ModelScore
1O3 (High)85.85
2O4 Mini (High)80.87
3Gemini 2.5 Pro (03-25)77.84
4Grok 3 Mini72.63
5Claude 3.7 Sonnet (Thinking)67.34
6Gemini 2.5 Flash (Thinking)63.69
7GPT-4.1 Mini54.55
8DeepSeek V3 (0324)50.91
9GPT-4.548.24
10Gemini 2.5 Flash47.61
11GPT-4.140.87
12Llama 4 Maverick34.74
13Llama 4 Scout16.16

Interactive version: theaggregate.ai/benchmark?slug=llm2014-logic-2025-04 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.