CUDAHercules - Single-Kernel Mean Speed (Self-Refine): leaderboard

Metric: Mean expert-relative speed (the expert reference runtime divided by the generated code's runtime, averaged over the 20 architecture-general Class 1 single-kernel tasks, a task whose code fails the validator counting 0, so 1 matches the expert; self-refine@10: up to 10 revision rounds with compile and execution feedback (CUDAForge)). Source: arxiv.org. Saturation forecast: Around July 2027. 4 models tracked.

Top models

#ModelScore
1GPT-5.40.31
2Claude Opus 4.60.27
3Qwen 3.5 122B A10B0.24

Interactive version: theaggregate.ai/benchmark?slug=cudahercules-single-kernel-mean-speed-self-refine · How It Works · Data refreshed daily, snapshot 2026-09-29.