Generalization V2 (Lechmazur) — leaderboard

Tests whether LLMs can infer a specific latent theme from limited examples, reject broader but incorrect patterns using anti-examples, and identify the single correct match among eight similar candidates. V2 uses 1,247 validated prompts.

Metric: Inverse-Rank Score. Source: github.com. Status: saturation imminent. 30 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Thinking, High)80.6
2GPT-5.4 (xHigh)80
3Gemini 3.1 Pro (Preview)79.4
4GPT-5.4 (High)77.4
5Claude Sonnet 4.6 (Thinking, High)76.3
6GLM-5.169.8
7Kimi K2.5 (Thinking)69.4
8Claude Opus 4.6 (Non-reasoning)68.8
9Qwen 3.5 397B A17B65.1
10DeepSeek V3.265
11Grok 4.20 0309 (Reasoning)63.8
12Gemini 3.1 Flash Lite (Preview)63.3
13GPT-5.4 Mini (xHigh)61.7
14Qwen 3.6 Plus59.5
15Seed 2.0 Pro57.1

Interactive version: theaggregate.ai/benchmark?slug=generalization-v2-lechmazur · How the rankings work · Data refreshed daily, snapshot 2026-07-22.