CESBench - Multiple Choice: leaderboard

Metric: Accuracy (%; the 209 four-option multiple-choice items whose distractors encode practitioner misconceptions; 380 expert-written items on cryptographic engineering security for IoT devices; zero-shot, temperature 0, no tools, provider default reasoning mode; judgment and scenario answers scored by a Qwen3.5-397B judge from outside the evaluated set). Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.

Top models

#ModelScore
1GLM-5.298.6
2Gemini 3.7 Flash98.6
3GPT-5.6 Luna97.6
4DeepSeek V4 Pro97.1
5Kimi K2.697.1
6MiniMax-M2.595.2
7Llama 4 Maverick93.3
8DeepSeek V391.9
9GLM-4 32B (0414)88.5
10Hunyuan A13B-Instruct85.6
11Ling-flash-2.083.3

Interactive version: theaggregate.ai/benchmark?slug=cesbench-multiple-choice · How It Works · Data refreshed daily, snapshot 2026-09-26.