SciCode — leaderboard
Scientific coding benchmark with 338 subproblems from 80 main problems across mathematics, physics, chemistry, biology, and materials science. Tests numerical implementation of research-level science.
Metric: Subproblem Resolve Rate (%). Source: scicode-bench.github.io. Status: saturation imminent. 20 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | O3 Mini (High) | 34.4 |
| 2 | O3 Mini (Low) | 33.3 |
| 3 | O3 Mini (Medium) | 33 |
| 4 | DeepSeek R1 | 28.5 |
| 5 | O1 Preview | 28.5 |
| 6 | Claude 3.5 Sonnet | 26 |
| 7 | GPT-4o | 25 |
| 8 | DeepSeek V3 | 23.7 |
| 9 | GPT-4 Turbo | 22.9 |
| 10 | O1 Mini | 22.2 |
| 11 | Gemini 1.5 Pro | 21.9 |
| 12 | Claude 3 Opus | 21.5 |
| 13 | DeepSeek Coder V2 | 21.2 |
| 14 | Claude 3 Sonnet | 17 |
| 15 | Qwen 2 72B Instruct | 17 |
Interactive version: theaggregate.ai/benchmark?slug=scicode · How the rankings work · Data refreshed daily, snapshot 2026-07-22.