KernelCraft (PLENA): leaderboard

Metric: Kernel success rate (%) on the PLENA accelerator: functionally correct kernels within the iteration budget (15, 20 and 25 iterations for primitive, composite and end-to-end tasks) over the 100 task configurations that apply to it, read from the Total successful/total count; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 4 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.2 (Medium)56#105 (GPT-5.2)
2Gemini 3 Flash (Preview) (Medium)35#78 (Gemini 3 Flash (Preview))
3Claude Sonnet 4 (20250514) (Thinking)12#211 (Claude Sonnet 4 (20250514))
4DeepSeek R1 05281#217

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=kernelcraft-plena · How It Works · Data refreshed daily, snapshot 2026-10-11.