Llama 2 7B Chat — benchmark results
Chat-tuned Llama 2 7B checkpoint. Provider: Meta. Released 2023-07-18. Access: Open.
Unified ELO 1273 ± 14, rank #1583 of 1776 rated models, from 278 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Trustworthy - Privacy | 97.39 | Trust Score (%) | 96 |
| LLM Trustworthy - Fairness | 100 | Trust Score (%) | 94 |
| JustEval - Safety | 5 | Score (1-5) | 93.3 |
| MMLU-by-task - Econometrics | 37.72 | Accuracy (%) | 89.6 |
| SALAD-Bench Base | 96.51 | Safety Score (%) | 84.8 |
| LLM Trustworthy - Out-of-Distribution | 75.65 | Trust Score (%) | 84 |
| LLM Trustworthy - Toxicity | 80 | Trust Score (%) | 84 |
| MMLU-by-task - College Mathematics | 36 | Accuracy (%) | 80.3 |
| LLM Trustworthy Leaderboard | 74.72 | Average Trust Score (%) | 80 |
| HumanLikeness - Sound-1 | 73.76 | Humanlike Score (%) | 78.9 |
| LLM Trustworthy - Adversarial | 51.01 | Trust Score (%) | 72 |
| LLM Trustworthy - Stereotype | 97.6 | Trust Score (%) | 72 |
Interactive version: theaggregate.ai/model?slug=llama-2-7b-chat · How the rankings work · Data refreshed daily, snapshot 2026-07-22.