Llama-3.1-Tulu-3-8B-DPO — benchmark results

Ai2's DPO-stage checkpoint of the fully open Tulu 3 post-training recipe, built on Llama 3.1 8B (November 2024). Provider: Meta. Released 2024-11-20. Access: Open.

Unified ELO 1495 ± 15, rank #819 of 1776 rated models, from 47 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard - IFEval80.29Score97.7
EuroEval Portuguese NLU - MultiWikiQA PT76.29Reading comprehension Score (%)90.1
EuroEval Dutch NLU - DBRD91.49Sentiment classification Score (%)87.3
EuroEval Italian NLU - ScaLA IT32.85Linguistic acceptability Score (%)80.6
EuroEval Finnish NLU - Scandisent FI91.52Sentiment classification Score (%)79
Open LLM Leaderboard - MATH Level 523.64Score77.1
EuroEval Spanish NLU - MLQA ES63.89Reading comprehension Score (%)74.8
EuroEval Spanish NLU - Sentiment Headlines ES45.55Sentiment classification Score (%)71.8
EuroEval Portuguese NLU54.33NLU Average Score (%)71.4
EuroEval Portuguese NLU - HAREM45.99Named entity recognition Score (%)65.9
EuroEval Italian NLU - SQuAD IT69.33Reading comprehension Score (%)64.9
EuroEval Spanish NLU46.59NLU Average Score (%)63.2

Interactive version: theaggregate.ai/model?slug=llama-3-1-tulu-3-8b-dpo · How the rankings work · Data refreshed daily, snapshot 2026-07-22.