CESBench - Implementation: leaderboard

Metric: Composite score (%; mean over the four task types of the item-score means within the D3 implementation sub-domain (83 items); 380 expert-written items on cryptographic engineering security for IoT devices; zero-shot, temperature 0, no tools, provider default reasoning mode; judgment and scenario answers scored by a Qwen3.5-397B judge from outside the evaluated set). Source: arxiv.org. Saturation forecast: Around December 2026. 11 models tracked.

Top models

#ModelScore
1GLM-5.285.4
2DeepSeek V4 Pro78.5
3Gemini 3.7 Flash78.5
4GPT-5.6 Luna77.9
5Kimi K2.675.6
6MiniMax-M2.569.9
7DeepSeek V362.5
8Ling-flash-2.061.4
9GLM-4 32B (0414)57.9
10Hunyuan A13B-Instruct54.8
11Llama 4 Maverick49.6

Interactive version: theaggregate.ai/benchmark?slug=cesbench-implementation · How It Works · Data refreshed daily, snapshot 2026-09-26.