WeirdML — leaderboard

17 novel, unusual ML tasks requiring genuine understanding - not memorization. Models generate PyTorch code in Docker containers with 5 iterative debug rounds. Tests shape recognition, chess prediction, and more.

Metric: Average Score. Source: htihle.github.io. Status: saturated. 139 models tracked.

Top models

#ModelScore
1Claude Fable 5 (Max)91.94
2GPT-5.6 Sol (High)88.76
3Claude Fable 5 (High)87.85
4GPT-5.5 (xHigh)84.91
5GPT-5.5 (High)83.9
6Claude Opus 4.8 (xHigh)82.89
7GPT-5.6 Terra (High)78.27
8Claude Opus 4.6 (High)77.95
9GPT-5.3 Codex (xHigh)77.9
10GPT-5.4 (xHigh)77.7
11Claude Opus 4.7 (High)76.44
12Claude Opus 4.7 (Max)75.45
13GPT-5.2 (xHigh)72.19
14Gemini 3.1 Pro (Preview) (High)72.07
15GLM-5.2 (Max)70.12

Interactive version: theaggregate.ai/benchmark?slug=weirdml · How the rankings work · Data refreshed daily, snapshot 2026-07-22.