Llama 3.1 8B: benchmark results

Meta's 8B pretrained base model from the Llama 3.1 family (July 2024), with 128K context and trained on over 15T tokens. Provider: Meta. Released 2024-07-23. Access: Open.

Unified ELO 1435 ± 1, rank #1013 of 1392 rated models, from 744 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SeaEval - Fundamental NLP Tasks - C3 (Few-Shot)81.04Accuracy (%)100
When Context Flips, Safety Breaks77.6PacifAIst BSR (self-reported)100
EuroEval Swedish NLU - Swerec80.44Sentiment classification Score (%)97.5
EuroEval German Summarization - Mlsum DE38.51Score (%)95.9
EuroEval Polish NLU - PoQuAD61.78Reading comprehension Score (%)95.7
EuroEval Lithuanian NLU - MultiWikiQA LT72.09Reading comprehension Score (%)95.6
EuroEval Greek NLU - MultiWikiQA EL72.93Reading comprehension Score (%)95.1
EuroEval Ukrainian NLU - MultiWikiQA UK66.87Reading comprehension Score (%)94.4
EuroEval Czech NLU - CSFD Sentiment64.67Sentiment classification Score (%)93.9
EuroEval Polish NLU - Polemo294.7Sentiment classification Score (%)93.6
LA Leaderboard - EusExams Basque44.9Accuracy (%)93.6
EuroEval Portuguese NLU - MultiWikiQA PT77.08Reading comprehension Score (%)92.9

Interactive version: theaggregate.ai/model?slug=llama-3-1-8b · How It Works · Data refreshed daily, snapshot 2026-09-05.