CyberBench (NLP) — leaderboard
Multi-task cybersecurity NLP evaluation covering named entity recognition, news classification, MITRE ATT&CK mapping, CVE analysis, and security Q&A across 9 subtasks.
Metric: Avg Score (%). Source: huggingface.co. Status: saturated. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4 | 74.03 |
| 2 | GPT-3.5 Turbo | 66.16 |
| 3 | Mistral-7B-v0.1 | 63.65 |
| 4 | zephyr-7B-beta | 61.07 |
| 5 | vicuna-13B-v1.5 | 60.26 |
| 6 | Llama 2 13B | 59.44 |
| 7 | Mistral 7B Instruct (v0.1) | 58.38 |
| 8 | Llama 2 7B | 55.61 |
| 9 | vicuna-7B-v1.5 | 55.52 |
| 10 | Llama 2 13B Chat | 49.2 |
| 11 | Llama 2 7B Chat | 47.16 |
| 12 | falcon-7B | 43.21 |
| 13 | falcon-7B Instruct | 40.76 |
Interactive version: theaggregate.ai/benchmark?slug=cyberbench-nlp · How the rankings work · Data refreshed daily, snapshot 2026-07-22.