LiveCodeBench Pro — leaderboard
Competitive coding problems at extreme difficulty. 'Hard' here means 99.9% of competitive coders can't solve them. Shows how LLMs stack up against elite human programmers.
Metric: Rating (CF-style). Source: livecodebenchpro.com. Status: years away from saturation. 58 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Deep Think | 3298 |
| 2 | Gemini 3.1 Pro (Preview) | 2887 |
| 3 | Gemini 3 Pro (Preview) | 2439 |
| 4 | Gemini 3 Flash (Preview) | 2316 |
| 5 | O4 Mini (2025-04-16) (High) | 2092 |
| 6 | Gemini 2.5 Pro | 1769 |
| 7 | O3 Mini (2025-01-31) | 1688 |
| 8 | Qwen 3 235B A22B (Thinking) | 1673 |
| 9 | Qwen 3 Next 80B A3B (Thinking) | 1603 |
| 10 | Claude Sonnet 4.5 (Thinking) | 1412 |
| 11 | Gemini 2.5 Flash (Preview 04-17) | 1334 |
| 12 | GPT-OSS-120B | 1299 |
| 13 | Gemini 2.5 Flash | 1288 |
| 14 | DeepSeek R1 0528 | 1284 |
| 15 | Qwen 3 Max | 1226 |
Interactive version: theaggregate.ai/benchmark?slug=livecodebench-pro · How the rankings work · Data refreshed daily, snapshot 2026-07-22.