CodeGolf Bench - Python Best Percentile: leaderboard

Metric: Best percentile (0-100): for each problem, the best of the sampled Python solutions placed in the distribution of human code.golf solutions by character count (0 = worse than every human or failing, 100 = shorter than every human), averaged over problems; 115 code.golf problems (holes), solutions checked against the platform's hidden tests through its API; golfing prompt, temperature 0.2, top-p 0.95, up to 32,768 output tokens, 10 samples per problem; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro (Preview 03-25)68.04
2DeepSeek R167.09
3Llama 4 Maverick55.31
4DeepSeek V3 (0324)51.6
5Gemini 2.5 Flash (Preview 04-17)46.42
6DeepSeek R1 Distill Llama 70B42.33
7Qwen 3 235B A22B41.64
8Llama 4 Scout36.69
9Gemma 3 27B (IT)25.23

Interactive version: theaggregate.ai/benchmark?slug=codegolf-bench-python-best-percentile · How It Works · Data refreshed daily, snapshot 2026-10-07.