Llama-3.1-Tulu-3-8B: benchmark results
Allen AI's fully open Tulu 3 post-train of Llama 3.1 8B, finishing with RLVR after SFT and DPO for strong math and instruction following (November 2024). Provider: Meta. Released 2024-11-20. Access: Open.
Unified ELO 1469 ± 1, rank #845 of 1392 rated models, from 65 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - IFEval | 82.67 | Score | 99.1 |
| EuroEval Dutch NLU - DBRD | 91.55 | Sentiment classification Score (%) | 87.8 |
| EuroEval Portuguese NLU - MultiWikiQA PT | 74.19 | Reading comprehension Score (%) | 80.5 |
| EuroEval Italian NLU - ScaLA IT | 32.25 | Linguistic acceptability Score (%) | 80.1 |
| Enkrypt AI - Jailbreak Risk | 4.34 | Risk Score | 78.8 |
| EuroEval Finnish NLU - Scandisent FI | 91.28 | Sentiment classification Score (%) | 76.5 |
| Open LLM Leaderboard - MATH Level 5 | 21.15 | Score | 72.9 |
| Enkrypt AI - Toxicity Risk | 1.82 | Risk Score | 72.3 |
| EuroEval Spanish NLU - MLQA ES | 63.29 | Reading comprehension Score (%) | 71.4 |
| HREF | 33.54 | Average HREF Score (%) | 69.7 |
| EuroEval Portuguese NLU | 53.53 | NLU Average Score (%) | 68.4 |
| EuroEval Spanish NLU - Sentiment Headlines ES | 44.51 | Sentiment classification Score (%) | 67.7 |
Interactive version: theaggregate.ai/model?slug=llama-3-1-tulu-3-8b · How It Works · Data refreshed daily, snapshot 2026-09-05.