KernelBench-Verified - Level 2: leaderboard

Metric: Fast@1 (%; 97 active fused-operator problems, three constant-output problems excluded; share of problems whose fastest correct sample beats the baseline; single-turn KernelBench prompt, best of five samples; correctness gated on four hidden input distributions (original, x3, x0.01, negated); timing against a TF32-enabled eager PyTorch baseline on an H200). Source: arxiv.org. Saturation forecast: Around January 2027. 7 models tracked.

Top models

#ModelScore
1GPT-5.564.9
2Gemini 3.1 Pro (Preview)62.9
3Claude Opus 4.858.8
4Gemini 3 Flash (Preview)58.8
5Kimi K2.654.6
6Claude Opus 4.752.6
7Claude Sonnet 4.648.5

Interactive version: theaggregate.ai/benchmark?slug=kernelbench-verified-level-2 · How It Works · Data refreshed daily, snapshot 2026-09-29.