KernelGenBench-MS (OpenCode) - vLLM Operators: leaderboard

Metric: Clean pass rate (%; share of the 50 vLLM inference operators whose generated Triton kernel passes numerical validation on every test case of its combinatorial shape, dtype and layout suite and every anti-hack check, as a drop-in replacement; the OpenCode agent with execution feedback; NVIDIA A100; 30-minute budget per operator). Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.

Top models

#ModelScore
1Claude Opus 4.658
2GLM-546
3Qwen 3.5 27B44
4MiniMax-M2.526

Interactive version: theaggregate.ai/benchmark?slug=kernelgenbench-ms-opencode-vllm-operators · How It Works · Data refreshed daily, snapshot 2026-09-29.