PerfCodeBench - Reference-Beating Rate: leaderboard

Metric: Reference-beating rate (%): share of comparable tasks where the correct code matches or beats the reference implementation, 1,854 system-level performance-optimization tasks across C, C++, Go, Java, Python and CUDA; the model returns one replacement source file from the task metadata, interface contract and baseline in one shot; a candidate earns performance credit only if it compiles, runs and passes the task oracle; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 20 models tracked.

Top models

#ModelScore
1GPT-561.6
2Kimi K2.652.13
3GPT-5.451.82
4DeepSeek V4 Pro47.95
5Claude Opus 4.543.23
6DeepSeek V4 Flash43.05
7Seed 2.0 Lite40.68
8Gemini 3.1 Pro (Preview)40.09
9Qwen 3.6 27B39.6
10Qwen 3.6 Max37.19
11Claude Sonnet 4.532.81
12Qwen 3.6 Plus29.7
13Kimi K227.58
14Qwen 3.6 35B A3B22.51
15Llama 4 Maverick20.3

Interactive version: theaggregate.ai/benchmark?slug=perfcodebench-reference-beating-rate · How It Works · Data refreshed daily, snapshot 2026-10-07.