Llama-3.1-Tulu-3-8B-DPO — benchmark results
Ai2's DPO-stage checkpoint of the fully open Tulu 3 post-training recipe, built on Llama 3.1 8B (November 2024). Provider: Meta. Released 2024-11-20. Access: Open.
Unified ELO 1495 ± 15, rank #819 of 1776 rated models, from 47 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - IFEval | 80.29 | Score | 97.7 |
| EuroEval Portuguese NLU - MultiWikiQA PT | 76.29 | Reading comprehension Score (%) | 90.1 |
| EuroEval Dutch NLU - DBRD | 91.49 | Sentiment classification Score (%) | 87.3 |
| EuroEval Italian NLU - ScaLA IT | 32.85 | Linguistic acceptability Score (%) | 80.6 |
| EuroEval Finnish NLU - Scandisent FI | 91.52 | Sentiment classification Score (%) | 79 |
| Open LLM Leaderboard - MATH Level 5 | 23.64 | Score | 77.1 |
| EuroEval Spanish NLU - MLQA ES | 63.89 | Reading comprehension Score (%) | 74.8 |
| EuroEval Spanish NLU - Sentiment Headlines ES | 45.55 | Sentiment classification Score (%) | 71.8 |
| EuroEval Portuguese NLU | 54.33 | NLU Average Score (%) | 71.4 |
| EuroEval Portuguese NLU - HAREM | 45.99 | Named entity recognition Score (%) | 65.9 |
| EuroEval Italian NLU - SQuAD IT | 69.33 | Reading comprehension Score (%) | 64.9 |
| EuroEval Spanish NLU | 46.59 | NLU Average Score (%) | 63.2 |
Interactive version: theaggregate.ai/model?slug=llama-3-1-tulu-3-8b-dpo · How the rankings work · Data refreshed daily, snapshot 2026-07-22.