CESBench - Code: leaderboard

Metric: Task success (%; the 41 Python and embedded C code tasks, success when every functional test passes and no process check (import allowlist, secret-independent control flow, Cortex-M0+ build and memory budget) fails; 380 expert-written items on cryptographic engineering security for IoT devices; zero-shot, temperature 0, no tools, provider default reasoning mode; judgment and scenario answers scored by a Qwen3.5-397B judge from outside the evaluated set). Source: arxiv.org. Saturation forecast: Around December 2026. 11 models tracked.

Top models

#ModelScore
1DeepSeek V4 Pro95.1
2Kimi K2.692.7
3GLM-5.292.7
4Gemini 3.7 Flash90.2
5GPT-5.6 Luna82.9
6MiniMax-M2.578
7DeepSeek V373.2
8Ling-flash-2.065.9
9Llama 4 Maverick53.7
10GLM-4 32B (0414)53.7
11Hunyuan A13B-Instruct41.5

Interactive version: theaggregate.ai/benchmark?slug=cesbench-code · How It Works · Data refreshed daily, snapshot 2026-09-26.