Llama 3.1 8B — benchmark results
Meta's 8B pretrained base model from the Llama 3.1 family (July 2024), with 128K context and trained on over 15T tokens. Provider: Meta. Released 2024-07-23. Access: Open.
Unified ELO 1401 ± 9, rank #1217 of 1776 rated models, from 721 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LIBRA - ruBABILongQA3 | 29.65 | Dataset Total Score (%) | 100 |
| SeaEval - Fundamental NLP Tasks - C3 (Few-Shot) | 81.04 | Accuracy (%) | 100 |
| When Context Flips, Safety Breaks | 77.6 | PacifAIst BSR (self-reported) | 100 |
| EuroEval Swedish NLU - Swerec | 80.44 | Sentiment classification Score (%) | 97.5 |
| MixRea | 20.3 | Implicit Information Ignorance Rate (I3R) (self-reported) | 97.5 |
| EuroEval German Summarization - Mlsum DE | 38.51 | Score (%) | 95.9 |
| EuroEval Polish NLU - PoQuAD | 61.78 | Reading comprehension Score (%) | 95.7 |
| EuroEval Lithuanian NLU - MultiWikiQA LT | 72.09 | Reading comprehension Score (%) | 95.6 |
| EuroEval Greek NLU - MultiWikiQA EL | 72.93 | Reading comprehension Score (%) | 95.1 |
| EuroEval Ukrainian NLU - MultiWikiQA UK | 66.87 | Reading comprehension Score (%) | 94.4 |
| EuroEval Czech NLU - CSFD Sentiment | 64.67 | Sentiment classification Score (%) | 93.9 |
| EuroEval Polish NLU - Polemo2 | 94.7 | Sentiment classification Score (%) | 93.6 |
Interactive version: theaggregate.ai/model?slug=llama-3-1-8b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.