LingOly-TOO — leaderboard

Linguistics reasoning benchmark evaluating models on baseline and obfuscated questions to separate reasoning ability from memorization.

Metric: Obfuscated Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 16 models tracked.

Top models

#ModelScore
1GPT-546.7
2Claude Opus 4.145.8
3Claude 3.7 Sonnet42.9
4Gemini 2.5 Pro42.4
5DeepSeek V3.1 Terminus42.2
6O1 Preview32.2
7O3 Mini (High)30.6
8Claude 3.5 Sonnet28.1
9GPT-4.525.5
10Gemini 1.5 Pro20.5
11GPT-4o15.6
12O3 Mini (Low)12.2
13Phi-411
14Llama 3.3 70B Instruct8.2
15aya-23-35B5.7

Interactive version: theaggregate.ai/benchmark?slug=lingoly-too · How the rankings work · Data refreshed daily, snapshot 2026-07-22.