OJBench — leaderboard
Competitive-programming benchmark drawn from NOI and ICPC-style online-judge tasks, evaluating algorithm design and code correctness.
Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 37 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro (03-25) | 38.91 |
| 2 | O4 Mini | 33.3 |
| 3 | O3 Mini | 31.79 |
| 4 | O1 | 26.45 |
| 5 | Qwen 3 235B A22B | 25.97 |
| 6 | DeepSeek V3 (0324) | 25.54 |
| 7 | O1 Mini | 21.61 |
| 8 | QwQ-32B | 19.02 |
| 9 | Claude 3.7 Sonnet (Thinking) | 18.27 |
| 10 | DeepSeek R1 Distill Llama 70B | 16.38 |
| 11 | Qwen 3 30B A3B | 15.84 |
| 12 | Qwen 3 32B | 14.92 |
| 13 | DeepSeek R1 Distill Qwen 14B | 14.17 |
| 14 | Gemini 2.0 Flash | 13.58 |
| 15 | Claude 3.5 Sonnet | 10.4 |
Interactive version: theaggregate.ai/benchmark?slug=ojbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.