Llama 3.1 8B — benchmark results

Meta's 8B pretrained base model from the Llama 3.1 family (July 2024), with 128K context and trained on over 15T tokens. Provider: Meta. Released 2024-07-23. Access: Open.

Unified ELO 1401 ± 9, rank #1217 of 1776 rated models, from 721 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LIBRA - ruBABILongQA329.65Dataset Total Score (%)100
SeaEval - Fundamental NLP Tasks - C3 (Few-Shot)81.04Accuracy (%)100
When Context Flips, Safety Breaks77.6PacifAIst BSR (self-reported)100
EuroEval Swedish NLU - Swerec80.44Sentiment classification Score (%)97.5
MixRea20.3Implicit Information Ignorance Rate (I3R) (self-reported)97.5
EuroEval German Summarization - Mlsum DE38.51Score (%)95.9
EuroEval Polish NLU - PoQuAD61.78Reading comprehension Score (%)95.7
EuroEval Lithuanian NLU - MultiWikiQA LT72.09Reading comprehension Score (%)95.6
EuroEval Greek NLU - MultiWikiQA EL72.93Reading comprehension Score (%)95.1
EuroEval Ukrainian NLU - MultiWikiQA UK66.87Reading comprehension Score (%)94.4
EuroEval Czech NLU - CSFD Sentiment64.67Sentiment classification Score (%)93.9
EuroEval Polish NLU - Polemo294.7Sentiment classification Score (%)93.6

Interactive version: theaggregate.ai/model?slug=llama-3-1-8b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.