CESBench: leaderboard

Metric: Composite score (%; unweighted mean of the four task-type means: multiple-choice accuracy (209), judgment verdict-gated rubric score (67), scenario rubric score (63) and code task success on 572 hidden tests (41); 380 expert-written items on cryptographic engineering security for IoT devices; zero-shot, temperature 0, no tools, provider default reasoning mode; judgment and scenario answers scored by a Qwen3.5-397B judge from outside the evaluated set). Source: arxiv.org. Saturation forecast: Around December 2026. 11 models tracked.

Top models

#ModelScore
1GLM-5.283.6
2Kimi K2.682.5
3Gemini 3.7 Flash81.4
4GPT-5.6 Luna80.8
5DeepSeek V4 Pro80.2
6MiniMax-M2.573.4
7DeepSeek V368.4
8Ling-flash-2.061.3
9GLM-4 32B (0414)60
10Llama 4 Maverick59.7
11Hunyuan A13B-Instruct54.4

Interactive version: theaggregate.ai/benchmark?slug=cesbench · How It Works · Data refreshed daily, snapshot 2026-09-26.