Llama-3.1-Tulu-3-8B: benchmark results

Allen AI's fully open Tulu 3 post-train of Llama 3.1 8B, finishing with RLVR after SFT and DPO for strong math and instruction following (November 2024). Provider: Meta. Released 2024-11-20. Access: Open.

Unified ELO 1469 ± 1, rank #845 of 1392 rated models, from 65 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard - IFEval82.67Score99.1
EuroEval Dutch NLU - DBRD91.55Sentiment classification Score (%)87.8
EuroEval Portuguese NLU - MultiWikiQA PT74.19Reading comprehension Score (%)80.5
EuroEval Italian NLU - ScaLA IT32.25Linguistic acceptability Score (%)80.1
Enkrypt AI - Jailbreak Risk4.34Risk Score78.8
EuroEval Finnish NLU - Scandisent FI91.28Sentiment classification Score (%)76.5
Open LLM Leaderboard - MATH Level 521.15Score72.9
Enkrypt AI - Toxicity Risk1.82Risk Score72.3
EuroEval Spanish NLU - MLQA ES63.29Reading comprehension Score (%)71.4
HREF33.54Average HREF Score (%)69.7
EuroEval Portuguese NLU53.53NLU Average Score (%)68.4
EuroEval Spanish NLU - Sentiment Headlines ES44.51Sentiment classification Score (%)67.7

Interactive version: theaggregate.ai/model?slug=llama-3-1-tulu-3-8b · How It Works · Data refreshed daily, snapshot 2026-09-05.