PerfCodeBench - Correct-and-Runnable Rate: leaderboard

Metric: Correct-and-runnable rate (%): share of tasks whose generated code compiles, runs and passes the oracle (compilation errors, runtime errors and wrong output all count as failures), 1,854 system-level performance-optimization tasks across C, C++, Go, Java, Python and CUDA; the model returns one replacement source file from the task metadata, interface contract and baseline in one shot; a candidate earns performance credit only if it compiles, runs and passes the task oracle; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 20 models tracked.

Top models

#ModelScore
1GPT-5.471.25
2Claude Opus 4.570.55
3Claude Sonnet 4.569.36
4GPT-561.81
5DeepSeek V4 Pro54.48
6Seed 2.0 Lite47.57
7Qwen 3.6 Max47.25
8DeepSeek V4 Flash47.09
9Qwen 3.6 Plus45.2
10Gemma 4 26B A4B45.15
11Gemma 4 31B42.45
12Kimi K236.03
13Qwen 3.6 35B A3B32.09
14Llama 4 Maverick29.4
15Gemini 3.1 Pro (Preview)25.57

Interactive version: theaggregate.ai/benchmark?slug=perfcodebench-correct-and-runnable-rate · How It Works · Data refreshed daily, snapshot 2026-10-07.