WeirdML: leaderboard

17 unusual ML tasks requiring genuine understanding - not memorization. Models generate PyTorch code in Docker containers with 5 iterative debug rounds. Tests shape recognition, chess prediction, and more.

Metric: Average Score. Source: htihle.github.io. Status: saturation imminent. 158 models tracked.

Top models

#ModelScore
1Claude Fable 5.1 (Max)92.9
2Claude Fable 5.1 (High)92.3
3Claude Fable 5 (Max)91.94
4Claude Opus 5 (Max)91.78
5Claude Opus 5 (High)91.59
6GPT-5.6 Sol (High)88.76
7Claude Fable 5 (High)87.85
8GPT-5.6 Sol (Max)86.97
9GPT-5.5 (xHigh)84.91
10GPT-5.5 (High)83.9
11Claude Opus 4.8 (xHigh)82.89
12Kimi K3 (Max)82.57
13GPT-5.6 Terra (High)78.27
14Claude Opus 4.6 (High)77.95
15GPT-5.3 Codex (xHigh)77.9

Interactive version: theaggregate.ai/benchmark?slug=weirdml · How It Works · Data refreshed daily, snapshot 2026-09-05.