Llama 2 7B Chat: benchmark results
Chat-tuned Llama 2 7B checkpoint. Provider: Meta. Released 2023-07-18. Access: Open.
Unified ELO 1360 ± 1, rank #1247 of 1392 rated models, from 305 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Trustworthy - Privacy | 97.39 | Trust Score (%) | 96 |
| LLM Trustworthy - Fairness | 100 | Trust Score (%) | 94 |
| JustEval - Safety | 5 | Score (1-5) | 93.3 |
| MMLU-by-task - Econometrics | 37.72 | Accuracy (%) | 89.6 |
| Enkrypt AI - Toxicity Risk | 0.5 | Risk Score | 88.7 |
| Enkrypt AI - Safety Risk | 19.64 | Risk Score | 85.9 |
| SALAD-Bench Base | 96.51 | Safety Score (%) | 84.8 |
| LLM Trustworthy - Out-of-Distribution | 75.65 | Trust Score (%) | 84 |
| LLM Trustworthy - Toxicity | 80 | Trust Score (%) | 84 |
| Enkrypt AI - Jailbreak Risk | 3.95 | Risk Score | 82.4 |
| MMLU-by-task - College Mathematics | 36 | Accuracy (%) | 80.3 |
| LLM Trustworthy Leaderboard | 74.72 | Average Trust Score (%) | 80 |
Interactive version: theaggregate.ai/model?slug=llama-2-7b-chat · How It Works · Data refreshed daily, snapshot 2026-09-05.