DataKernelBench: leaderboard

Metric: Overall speedup over TorchPlan with torch.compile (times faster, end-to-end run_query runtime; a query whose kernel is incorrect or under the minimum speedup is credited at the torch.compile runtime; 22 TPC-H queries at scale factor 10 on one NVIDIA H100, each translated into a validated PyTorch TorchPlan; the model rewrites the query core or the full query in CUDA or Triton with up to several rounds of execution-guided repair; each model's best of its four framework-level configurations, as the leaderboard reports). Source: arxiv.org. Saturation forecast: Around June 2027. 10 models tracked.

Top models

#ModelScore
1GPT-5.52.11
2Claude Sonnet 4.61.54
3Claude Opus 4.71.51
4Gemini 3.1 Pro (Preview)1.44
5Claude Haiku 4.51.3
6GPT-OSS-120B1.26
7Qwen 3.5 397B A17B1.26
8DeepSeek V4 Flash1.23
9MiniMax-M2.51.19
10Devstral 21.08

Interactive version: theaggregate.ai/benchmark?slug=datakernelbench · How It Works · Data refreshed daily, snapshot 2026-09-29.