CUDABeaver - Pass@1: leaderboard

Metric: Share of the 213 broken-start CUDA debugging tasks (%) whose first repair attempt compiles, passes the task tests and runs at least 0.7 times as fast as the reference (performance gate p=0.7), default protocol: category-aware error feedback (level L3), iterative sampling at temperature 0.7; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 6 models tracked.

Top models

#ModelScore
1Qwen 3.6 Plus15.5
2Gemma 4 31B (IT)14.7
3GPT-5.414.5
4Kimi K2.613.6
5Qwen 3.6 27B10.3
6GLM-4.77.5

Interactive version: theaggregate.ai/benchmark?slug=cudabeaver-pass-1 · How It Works · Data refreshed daily, snapshot 2026-10-07.