CESBench - Scenario: leaderboard

Metric: Scenario score (%; the 63 open-ended engineering diagnosis items, mean of 3 to 5 expert rubric dimensions judged 0, 0.5 or 1; 380 expert-written items on cryptographic engineering security for IoT devices; zero-shot, temperature 0, no tools, provider default reasoning mode; judgment and scenario answers scored by a Qwen3.5-397B judge from outside the evaluated set). Source: arxiv.org. Saturation forecast: Around December 2026. 11 models tracked.

Top models

#ModelScore
1GPT-5.6 Luna88.4
2GLM-5.284.2
3Kimi K2.682.2
4Gemini 3.7 Flash81.9
5DeepSeek V4 Pro81.4
6MiniMax-M2.571.9
7DeepSeek V367.5
8GLM-4 32B (0414)59.1
9Llama 4 Maverick57.1
10Ling-flash-2.055.9
11Hunyuan A13B-Instruct47.2

Interactive version: theaggregate.ai/benchmark?slug=cesbench-scenario · How It Works · Data refreshed daily, snapshot 2026-09-26.