KernelGenBench-MS (Pass@1) - vLLM Operators: leaderboard

Metric: Clean pass rate (%; share of the 50 vLLM inference operators whose generated Triton kernel passes numerical validation on every test case of its combinatorial shape, dtype and layout suite and every anti-hack check, as a drop-in replacement; direct zero-shot generation, one sample at temperature 0, no execution feedback; NVIDIA A100; 30-minute budget per operator). Source: arxiv.org. Saturation forecast: Around August 2027. 4 models tracked.

Top models

#ModelScore
1GLM-524
2Claude Opus 4.620
3Qwen 3.5 27B2
4MiniMax-M2.50

Interactive version: theaggregate.ai/benchmark?slug=kernelgenbench-ms-pass-1-vllm-operators · How It Works · Data refreshed daily, snapshot 2026-09-29.