SciCodePile: leaderboard

Metric: Pass@1 (%) on the 200 executable scientific Python function-generation tasks, each run in a sandbox with stubbed dependencies against its test harness (HumanEval protocol); higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 15 models tracked.

Top models

#ModelScore
1GPT-5.4 Mini12.3
2O3 Mini10.5
3DeepSeek R18.5
4Qwen 3 14B (Reasoning)8.3
5GPT-4o7.5
6Qwen 2.5 Coder 7B Instruct6.1
7GPT-55.7
8DeepSeek R1 Distill Qwen 1.5B2.7
9starcoder2-3B0.5
10starcoder2-7B0.3
11starcoder2-15B0.3

Interactive version: theaggregate.ai/benchmark?slug=scicodepile · How It Works · Data refreshed daily, snapshot 2026-09-29.