MMLU-by-task - Security Studies — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 1257 models tracked.

Top models

#ModelScore
1Llama 2 70B Base79.18
2Llama 2 70B Chat (HF)78.78
3falcon-180B78.37
4Llama 2 70B Chat76.73
5Llama 2 70B Chat GPTQ75.92
6StableBeluga275.51
7Mistral-7B-v0.172.65
8LLaMA-65B71.43
9vicuna-33B-v1.368.98
10internlm-20B Chat68.98
11LLaMA-30B66.94
12WizardLM-13B-V1.266.94
13OpenHermes-13B64.9
14trurl-2-13B-academic64.49
15Llama 2 13B Chat Base64.08

Interactive version: theaggregate.ai/benchmark?slug=mmlu-by-task-security-studies · How the rankings work · Data refreshed daily, snapshot 2026-07-22.