Generalization V1 (Lechmazur) — leaderboard

Tests whether LLMs can infer a specific latent theme from limited examples, reject broader but incorrect patterns using anti-examples, and identify the single correct match among similar candidates. V1 uses 810 prompts.

Metric: Avg Rank (lower is better). Source: github.com. Status: saturated. 82 models tracked.

Top models

#ModelScore
1Claude Opus 4 (Thinking 16K)1.69
2Claude Opus 4.1 (Non-reasoning)1.69
3Claude Sonnet 4 (Thinking 16K)1.69
4Claude Opus 4 (Non-reasoning)1.7
5Gemini 2.5 Pro1.74
6DeepSeek R1 05281.74
7Gemini 2.5 Pro (Preview 05-06)1.75
8GPT-5 (Medium)1.79
9Grok 4 Fast (Reasoning)1.79
10DeepSeek R11.8
11O11.8
12GLM-4.51.8
13O4 Mini (Medium)1.8
14O4 Mini (High)1.82
15O3 (High)1.82

Interactive version: theaggregate.ai/benchmark?slug=generalization-v1-lechmazur · How the rankings work · Data refreshed daily, snapshot 2026-07-22.