Llama 2 13B Chat Base — benchmark results
Provider: Meta. Released 2023-07-18. Access: Open.
Unified ELO 1336 ± 15, rank #1452 of 1776 rated models, from 175 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HumanLikeness - Discourse-2 | 78.85 | Humanlike Score (%) | 100 |
| Open Japanese LLM - Xlsum JA Bleu JA | 33.8 | Score (%) | 95.4 |
| HumanLikeness - Word-2 | 25.05 | Humanlike Score (%) | 94.7 |
| SALAD-Bench | 81.27 | Average Safety Score (%) | 93.9 |
| SALAD-Bench Base | 96.81 | Safety Score (%) | 89.4 |
| MMLU-by-task - High School Chemistry | 46.31 | Accuracy (%) | 88.5 |
| SALAD-Bench Attack | 65.72 | Safety Score (%) | 87.9 |
| MMLU-by-task - International Law | 76.86 | Accuracy (%) | 87.1 |
| MMLU-by-task - Public Relations | 66.36 | Accuracy (%) | 85.5 |
| HumanLikeness - Meaning-1 | 73.9 | Humanlike Score (%) | 84.2 |
| MMLU-by-task - Virology | 48.19 | Accuracy (%) | 83.9 |
| MMLU-by-task - High School Mathematics | 31.11 | Accuracy (%) | 83.3 |
Interactive version: theaggregate.ai/model?slug=llama-2-13b-chat-base · How the rankings work · Data refreshed daily, snapshot 2026-07-22.