DataKernelBench - Pass Rate: leaderboard

Metric: Pass rate (%; share of the 22 queries whose optimized module matches the reference output, for the configuration the leaderboard reports; 22 TPC-H queries at scale factor 10 on one NVIDIA H100, each translated into a validated PyTorch TorchPlan; the model rewrites the query core or the full query in CUDA or Triton with up to several rounds of execution-guided repair; each model's best of its four framework-level configurations, as the leaderboard reports). Source: arxiv.org. Saturation forecast: Estimated already saturated. 10 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.6100
2GPT-5.5100
3Claude Opus 4.7100
4Qwen 3.5 397B A17B100
5Claude Haiku 4.595.5
6DeepSeek V4 Flash90.9
7GPT-OSS-120B86.4
8MiniMax-M2.581.8
9Gemini 3.1 Pro (Preview)54.5
10Devstral 245.5

Interactive version: theaggregate.ai/benchmark?slug=datakernelbench-pass-rate · How It Works · Data refreshed daily, snapshot 2026-09-29.