OJBench — leaderboard

Competitive-programming benchmark drawn from NOI and ICPC-style online-judge tasks, evaluating algorithm design and code correctness.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 37 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro (03-25)38.91
2O4 Mini33.3
3O3 Mini31.79
4O126.45
5Qwen 3 235B A22B25.97
6DeepSeek V3 (0324)25.54
7O1 Mini21.61
8QwQ-32B19.02
9Claude 3.7 Sonnet (Thinking)18.27
10DeepSeek R1 Distill Llama 70B16.38
11Qwen 3 30B A3B15.84
12Qwen 3 32B14.92
13DeepSeek R1 Distill Qwen 14B14.17
14Gemini 2.0 Flash13.58
15Claude 3.5 Sonnet10.4

Interactive version: theaggregate.ai/benchmark?slug=ojbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.