LingOly-TOO: leaderboard

Linguistics reasoning benchmark evaluating models on baseline and obfuscated questions to separate reasoning ability from memorization.

Metric: Obfuscated Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 16 models tracked.

Top models

#ModelScore
1GPT-546.7
2Claude Opus 4.145.8
3Claude 3.7 Sonnet42.89
4Gemini 2.5 Pro42.35
5DeepSeek V3.1 Terminus42.2
6O1 Preview32.22
7O3 Mini (High)30.59
8Claude 3.5 Sonnet28.1
9DeepSeek R126.5
10GPT-4.525.45
11Gemini 1.5 Pro20.46
12GPT-4o15.63
13O3 Mini (Low)12.22
14Phi-411
15Llama 3.3 70B Instruct8.21

Interactive version: theaggregate.ai/benchmark?slug=lingoly-too · How It Works · Data refreshed daily, snapshot 2026-09-05.