CyberMetric — leaderboard

Multiple-choice cybersecurity knowledge Q&A across 500 questions covering threats, vulnerabilities, defenses, and best practices. Tests breadth of security knowledge.

Metric: Accuracy (%). Source: huggingface.co. Status: saturated. 25 models tracked.

Top models

#ModelScore
1GPT-4o92.45
2GPT-4 Turbo92.26
3Mixtral 8x7B Instruct90.6
4Gemini 1.0 Pro86.64
5GPT-3.5 Turbo86.42
6Mistral 7B Instruct (v0.2)77.09
7Llama 3 8B Instruct75.61
8zephyr-7B-beta73.71
9flan-t5-xxl72.38
10Llama 2 70B71.53
11Qwen 1.5 7B63.87
12Qwen1.5-MoE-A2.7B62.37
13Qwen-7B52.9
14Phi-251.7
15Llama3-ChatQA-1.5-8B51.41

Interactive version: theaggregate.ai/benchmark?slug=cybermetric · How the rankings work · Data refreshed daily, snapshot 2026-07-22.