PerfCodeBench - Efficiency Gap Closed: leaderboard

Metric: Share of comparable tasks (%) whose correct code closes at least 80 percent of the baseline-to-reference performance gap (CGRE at least 0.8), 1,854 system-level performance-optimization tasks across C, C++, Go, Java, Python and CUDA; the model returns one replacement source file from the task metadata, interface contract and baseline in one shot; a candidate earns performance credit only if it compiles, runs and passes the task oracle; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 20 models tracked.

Top models

#ModelScore
1GPT-572.91
2GPT-5.464.26
3Claude Opus 4.563.17
4Qwen 3.6 Max63.1
5DeepSeek V4 Pro62.88
6Kimi K2.661.7
7Qwen 3.6 27B61.39
8Gemini 3.1 Pro (Preview)60.67
9DeepSeek V4 Flash60.29
10Claude Sonnet 4.559.15
11Seed 2.0 Lite58.99
12Qwen 3.6 Plus54.85
13Kimi K243.58
14Qwen 3.6 35B A3B41.26
15Llama 4 Maverick40.26

Interactive version: theaggregate.ai/benchmark?slug=perfcodebench-efficiency-gap-closed · How It Works · Data refreshed daily, snapshot 2026-10-07.