Llama 2 13B Chat Base — benchmark results

Provider: Meta. Released 2023-07-18. Access: Open.

Unified ELO 1336 ± 15, rank #1452 of 1776 rated models, from 175 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HumanLikeness - Discourse-278.85Humanlike Score (%)100
Open Japanese LLM - Xlsum JA Bleu JA33.8Score (%)95.4
HumanLikeness - Word-225.05Humanlike Score (%)94.7
SALAD-Bench81.27Average Safety Score (%)93.9
SALAD-Bench Base96.81Safety Score (%)89.4
MMLU-by-task - High School Chemistry46.31Accuracy (%)88.5
SALAD-Bench Attack65.72Safety Score (%)87.9
MMLU-by-task - International Law76.86Accuracy (%)87.1
MMLU-by-task - Public Relations66.36Accuracy (%)85.5
HumanLikeness - Meaning-173.9Humanlike Score (%)84.2
MMLU-by-task - Virology48.19Accuracy (%)83.9
MMLU-by-task - High School Mathematics31.11Accuracy (%)83.3

Interactive version: theaggregate.ai/model?slug=llama-2-13b-chat-base · How the rankings work · Data refreshed daily, snapshot 2026-07-22.