LiveCodeBench Pro: leaderboard
Competitive coding problems at extreme difficulty. 'Hard' here means 99.9% of competitive coders can't solve them. Shows how LLMs stack up against elite human programmers.
Metric: Rating (CF-style). Source: livecodebenchpro.com. Status: saturation imminent. 58 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Deep Think | 3298 |
| 2 | Gemini 3.1 Pro (Preview) | 2887 |
| 3 | Gemini 3 Pro (Preview) | 2439 |
| 4 | GPT-5.2 (High) | 2393 |
| 5 | Gemini 3 Flash (Preview) | 2316 |
| 6 | GPT-5 (High) | 2176 |
| 7 | O4 Mini (2025-04-16) (High) | 2092 |
| 8 | Gemini 2.5 Pro | 1769 |
| 9 | Gemini 2.5 Pro (Preview 03-25) | 1694 |
| 10 | O3 Mini (2025-01-31) | 1688 |
| 11 | Qwen 3 235B A22B 2507 (Thinking) | 1673 |
| 12 | Qwen 3 Next 80B A3B (Thinking) | 1603 |
| 13 | Claude Sonnet 4.5 (Thinking) | 1412 |
| 14 | Gemini 2.5 Flash (Preview 04-17) | 1334 |
| 15 | GPT-OSS-120B | 1299 |
Interactive version: theaggregate.ai/benchmark?slug=livecodebench-pro · How It Works · Data refreshed daily, snapshot 2026-09-05.