Llama 3.1 8B Instruct — benchmark results
Meta Llama 3.1 8B instruction-tuned checkpoint. Provider: Meta. Released 2024-07-23. Access: Open.
Unified ELO 1437 ± 5, rank #1057 of 1776 rated models, from 1071 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AgentCollabBench | 10.1 | IDR (self-reported) | 100 |
| LIBRA - MatreshkaYesNo | 75 | Dataset Total Score (%) | 100 |
| MASLegalBench | 86.21 | Best accuracy (self-reported) | 100 |
| Open CoT - LSAT Reading Comprehension | 27.51 | CoT Gain (%) | 100 |
| EuroEval Serbian Summarization - LR SUM SR | 31.53 | Score (%) | 99.3 |
| Open CoT Leaderboard | 16.29 | Average CoT Gain (%) | 98.5 |
| EuroEval Bosnian Summarization - LR SUM BS | 32.19 | Score (%) | 98 |
| EuroEval Dutch Summarization - Wiki Lingua NL | 36.04 | Score (%) | 98 |
| EVALITA - evalita NER | 40.3 | CPS | 97.9 |
| EuroEval Albanian Summarization - LR SUM SQ | 34.7 | Score (%) | 97.4 |
| EuroEval Catalan Summarization - Dacsa CA | 39.33 | Score (%) | 97.4 |
| LatamBoard - Spanish PAWS | 64.85 | Score (%) | 97 |
Interactive version: theaggregate.ai/model?slug=llama-3-1-8b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.