PerfCodeBench - Faster-than-Baseline Rate: leaderboard

Metric: Faster-than-baseline rate (%): share of tasks where the model produces correct code that runs faster than the baseline implementation, 1,854 system-level performance-optimization tasks across C, C++, Go, Java, Python and CUDA; the model returns one replacement source file from the task metadata, interface contract and baseline in one shot; a candidate earns performance credit only if it compiles, runs and passes the task oracle; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 20 models tracked.

Top models

#ModelScore
1GPT-5.466.72
2Claude Opus 4.565.75
3Claude Sonnet 4.562.08
4GPT-559.55
5DeepSeek V4 Pro51.13
6DeepSeek V4 Flash44.82
7Seed 2.0 Lite43.8
8Qwen 3.6 Max42.5
9Qwen 3.6 Plus39.75
10Gemma 4 26B A4B39.64
11Gemma 4 31B37.65
12Kimi K232.9
13Qwen 3.6 35B A3B27.94
14Llama 4 Maverick24.38
15Gemini 3.1 Pro (Preview)23.73

Interactive version: theaggregate.ai/benchmark?slug=perfcodebench-faster-than-baseline-rate · How It Works · Data refreshed daily, snapshot 2026-10-07.