Llama-3.1-Tulu-3-8B — benchmark results

Allen AI's fully open Tulu 3 post-train of Llama 3.1 8B, finishing with RLVR after SFT and DPO for strong math and instruction following (November 2024). Provider: Meta. Released 2024-11-20. Access: Open.

Unified ELO 1492 ± 14, rank #830 of 1776 rated models, from 49 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard - IFEval82.67Score99.1
EuroEval Dutch NLU - DBRD91.55Sentiment classification Score (%)87.8
EuroEval Portuguese NLU - MultiWikiQA PT74.19Reading comprehension Score (%)80.5
EuroEval Italian NLU - ScaLA IT32.25Linguistic acceptability Score (%)80.1
EuroEval Finnish NLU - Scandisent FI91.28Sentiment classification Score (%)76.5
Open LLM Leaderboard - MATH Level 521.15Score72.9
EuroEval Spanish NLU - MLQA ES63.29Reading comprehension Score (%)71.4
HREF33.54Average HREF Score (%)69.7
EuroEval Portuguese NLU53.53NLU Average Score (%)68.4
EuroEval Spanish NLU - Sentiment Headlines ES44.51Sentiment classification Score (%)67.7
EuroEval Spanish NLU46.47NLU Average Score (%)62.9
EuroEval Portuguese NLU - HAREM44.94Named entity recognition Score (%)61.4

Interactive version: theaggregate.ai/model?slug=llama-3-1-tulu-3-8b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.