LiveOIBench — leaderboard

Contamination-free Olympiad benchmark with 403 expert-curated problems from 14 Informatics Olympiads (IOI, USACO, APIO, etc). Models ranked by human percentile against actual contestants.

Metric: Avg Human Percentile. Source: liveoibench.github.io. Status: saturation imminent. 58 models tracked.

Top models

#ModelScore
1GPT-580.18
2GPT-OSS-120B (High)71.85
3Grok 4 Fast (Reasoning)69.8
4Gemini 2.5 Pro67.44
5O3 Mini (High)60.86
6GPT-OSS-120B (Medium)57.55
7GPT-OSS-20B (High)57.13
8Gemini 2.5 Flash53.75
9Claude Sonnet 4.551.39
10GPT-OSS-20B (Medium)51.12
11GPT-OSS-120B (Low)44.07
12Qwen 3 32B41.95
13DeepSeek R138.8
14Qwen 3 30B A3B35.2
15GPT-OSS-20B (Low)34.64

Interactive version: theaggregate.ai/benchmark?slug=liveoibench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.