Llama 2 7B Chat — benchmark results

Chat-tuned Llama 2 7B checkpoint. Provider: Meta. Released 2023-07-18. Access: Open.

Unified ELO 1273 ± 14, rank #1583 of 1776 rated models, from 278 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Trustworthy - Privacy97.39Trust Score (%)96
LLM Trustworthy - Fairness100Trust Score (%)94
JustEval - Safety5Score (1-5)93.3
MMLU-by-task - Econometrics37.72Accuracy (%)89.6
SALAD-Bench Base96.51Safety Score (%)84.8
LLM Trustworthy - Out-of-Distribution75.65Trust Score (%)84
LLM Trustworthy - Toxicity80Trust Score (%)84
MMLU-by-task - College Mathematics36Accuracy (%)80.3
LLM Trustworthy Leaderboard74.72Average Trust Score (%)80
HumanLikeness - Sound-173.76Humanlike Score (%)78.9
LLM Trustworthy - Adversarial51.01Trust Score (%)72
LLM Trustworthy - Stereotype97.6Trust Score (%)72

Interactive version: theaggregate.ai/model?slug=llama-2-7b-chat · How the rankings work · Data refreshed daily, snapshot 2026-07-22.