CyberCertBench - Fortinet ICS/SCADA: leaderboard

Metric: Accuracy (%) on 43 Fortinet ICS/SCADA questions (conceptual operational-technology security): 5-shot multiple-choice question answering at temperature 0, scored by exact accuracy (correct only when the model selects all correct options and no incorrect one), mean accuracy (%) as printed in the appendix; questions scraped from public certification question banks and checked against authoritative documentation; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 45 models tracked.

Top models

#ModelScore
1GPT-588.4
2GPT-OSS-120B88.4
3Mistral Small 3.188.4
4Qwen 2.5 72B87.4
5GPT-4o86.5
6GPT-4.186
7Mistral Large 2 (Nov) Instruct (2411)86
8Gemini 2.5 Pro (Preview 05-06)86
9Qwen 2 72B86
10GPT-5 Mini83.7
11Claude Haiku 4.583.7
12Kimi K283.7
13Claude 3.5 Sonnet83.7
14Qwen 3 32B81.4
15Gemini 2.5 Flash Lite81.4

Interactive version: theaggregate.ai/benchmark?slug=cybercertbench-fortinet-ics-scada · How It Works · Data refreshed daily, snapshot 2026-10-07.