ALE-Bench — leaderboard

Score-based algorithmic programming benchmark built from AtCoder Heuristic Contest tasks, evaluating AI systems on hard optimization problems with hidden/private test evaluation.

Metric: Performance (Self-Refine x1) (self-reported). Source: benchmarklist.com. Status: saturation imminent. 79 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)1942.97
2GPT-5.3 Codex (xHigh)1655.22
3GPT-5.4 (High)1607
4Gemini 3 Flash (Preview) (High)1367.2
5Claude Opus 4.7 (Thinking)1323.05
6GPT-5.2 Codex (xHigh)1299.9
7GPT-5.2 (High)1293.55
8GPT-5.1 Codex (High)1244.92
9GPT-5.11192.15
10GPT-5.4 Mini (High)1188.58
11Gemini 3 Pro (High)1176.75
12GPT-51162.45
13Gemini 3.1 Pro (Preview) (High)1160.6
14Grok 4.201150.28
15Claude Opus 4.51025.38

Interactive version: theaggregate.ai/benchmark?slug=ale-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.