CESBench - Judgment: leaderboard

Metric: Judgment score (%; the 67 true-or-false security claims: a wrong verdict scores 0, a right one earns the mean of four justification rubric points judged 0, 0.5 or 1; 380 expert-written items on cryptographic engineering security for IoT devices; zero-shot, temperature 0, no tools, provider default reasoning mode; judgment and scenario answers scored by a Qwen3.5-397B judge from outside the evaluated set). Source: arxiv.org. Saturation forecast: Around 2029. 11 models tracked.

Top models

#ModelScore
1GLM-5.258.8
2Kimi K2.658
3Gemini 3.7 Flash54.9
4GPT-5.6 Luna54.5
5MiniMax-M2.548.3
6DeepSeek V4 Pro47
7Hunyuan A13B-Instruct43.3
8DeepSeek V341.2
9Ling-flash-2.040.1
10GLM-4 32B (0414)38.6
11Llama 4 Maverick34.9

Interactive version: theaggregate.ai/benchmark?slug=cesbench-judgment · How It Works · Data refreshed daily, snapshot 2026-09-26.