CUDAHercules - Single-Kernel Mean Speed (One-Shot): leaderboard
Metric: Mean expert-relative speed (the expert reference runtime divided by the generated code's runtime, averaged over the 20 architecture-general Class 1 single-kernel tasks, a task whose code fails the validator counting 0, so 1 matches the expert; one-shot: task specification, hardware metadata and few-shot exemplars, one sample, no execution feedback). Source: arxiv.org. Saturation forecast: Around July 2027. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 0.23 |
| 2 | Claude Opus 4.6 | 0.11 |
| 3 | Qwen 3.5 122B A10B | 0.06 |
Interactive version: theaggregate.ai/benchmark?slug=cudahercules-single-kernel-mean-speed-one-shot · How It Works · Data refreshed daily, snapshot 2026-09-29.