SciCode — leaderboard

Scientific coding benchmark with 338 subproblems from 80 main problems across mathematics, physics, chemistry, biology, and materials science. Tests numerical implementation of research-level science.

Metric: Subproblem Resolve Rate (%). Source: scicode-bench.github.io. Status: saturation imminent. 20 models tracked.

Top models

#ModelScore
1O3 Mini (High)34.4
2O3 Mini (Low)33.3
3O3 Mini (Medium)33
4DeepSeek R128.5
5O1 Preview28.5
6Claude 3.5 Sonnet26
7GPT-4o25
8DeepSeek V323.7
9GPT-4 Turbo22.9
10O1 Mini22.2
11Gemini 1.5 Pro21.9
12Claude 3 Opus21.5
13DeepSeek Coder V221.2
14Claude 3 Sonnet17
15Qwen 2 72B Instruct17

Interactive version: theaggregate.ai/benchmark?slug=scicode · How the rankings work · Data refreshed daily, snapshot 2026-07-22.