KernelBench Hub - Hard: leaderboard
KernelBench Hard evaluates autonomous coding agents on GPU kernel engineering tasks, measuring correctness and speed relative to hardware baselines.
Metric: Best % of Hardware Roofline. Source: kernelbench.com. Status: years away from saturation. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3.8 Max | 69.26 |
| 2 | GLM-5.3 | 63.19 |
| 3 | Grok 4.6 | 62.17 |
| 4 | GPT-5.6 Sol | 56.55 |
| 5 | DeepSeek V4 Flash (0731) | 48.6 |
| 6 | Kimi K3 | 48.55 |
| 7 | Claude Fable 5 | 43.03 |
| 8 | Claude Opus 5 | 37.08 |
| 9 | GLM-5.3 Flash | 10.21 |
Interactive version: theaggregate.ai/benchmark?slug=kernelbench-hub-hard · How It Works · Data refreshed daily, snapshot 2026-09-05.