CyberBench (NLP) — leaderboard

Multi-task cybersecurity NLP evaluation covering named entity recognition, news classification, MITRE ATT&CK mapping, CVE analysis, and security Q&A across 9 subtasks.

Metric: Avg Score (%). Source: huggingface.co. Status: saturated. 13 models tracked.

Top models

#ModelScore
1GPT-474.03
2GPT-3.5 Turbo66.16
3Mistral-7B-v0.163.65
4zephyr-7B-beta61.07
5vicuna-13B-v1.560.26
6Llama 2 13B59.44
7Mistral 7B Instruct (v0.1)58.38
8Llama 2 7B55.61
9vicuna-7B-v1.555.52
10Llama 2 13B Chat49.2
11Llama 2 7B Chat47.16
12falcon-7B43.21
13falcon-7B Instruct40.76

Interactive version: theaggregate.ai/benchmark?slug=cyberbench-nlp · How the rankings work · Data refreshed daily, snapshot 2026-07-22.