SciCodePile: leaderboard
Metric: Pass@1 (%) on the 200 executable scientific Python function-generation tasks, each run in a sandbox with stubbed dependencies against its test harness (HumanEval protocol); higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 Mini | 12.3 |
| 2 | O3 Mini | 10.5 |
| 3 | DeepSeek R1 | 8.5 |
| 4 | Qwen 3 14B (Reasoning) | 8.3 |
| 5 | GPT-4o | 7.5 |
| 6 | Qwen 2.5 Coder 7B Instruct | 6.1 |
| 7 | GPT-5 | 5.7 |
| 8 | DeepSeek R1 Distill Qwen 1.5B | 2.7 |
| 9 | starcoder2-3B | 0.5 |
| 10 | starcoder2-7B | 0.3 |
| 11 | starcoder2-15B | 0.3 |
Interactive version: theaggregate.ai/benchmark?slug=scicodepile · How It Works · Data refreshed daily, snapshot 2026-09-29.