CESBench - Evaluation: leaderboard

Metric: Composite score (%; mean over the four task types of the item-score means within the D5 evaluation and certification sub-domain (51 items); 380 expert-written items on cryptographic engineering security for IoT devices; zero-shot, temperature 0, no tools, provider default reasoning mode; judgment and scenario answers scored by a Qwen3.5-397B judge from outside the evaluated set). Source: arxiv.org. Saturation forecast: Around December 2026. 11 models tracked.

Top models

#ModelScore
1Kimi K2.680.5
2Gemini 3.7 Flash79.7
3GPT-5.6 Luna77.5
4DeepSeek V4 Pro77.3
5MiniMax-M2.573.6
6GLM-5.271.8
7DeepSeek V368.3
8Llama 4 Maverick60.2
9GLM-4 32B (0414)59.4
10Ling-flash-2.058.8
11Hunyuan A13B-Instruct53.4

Interactive version: theaggregate.ai/benchmark?slug=cesbench-evaluation · How It Works · Data refreshed daily, snapshot 2026-09-26.