Llama-3.1-Tulu-3-8B-SFT: benchmark results

Provider: Meta. Released 2024-11-21. Access: Open.

Unified ELO 1482 ± 20, rank #1255 of 2656 rated models, from 396 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EuroEval Ukrainian NLU - Cross Domain UK Reviews64.17Sentiment classification Score (%)98.5
AbstentionBench - subjective - CoCoNot/Humanizing - F198.77Abstention F1 (%)97.4
AbstentionBench - subjective - CoCoNot/Humanizing - Recall97.56Abstention Recall (%)97.4
EuroEval Icelandic NLU - NQII60.51Reading comprehension Score (%)94.8
AbstentionBench - answer unknown - CoCoNot/Unsupported - Recall96.67Abstention Recall (%)94.7
AbstentionBench - subjective - KUQ/Controversial - Precision96.15Abstention Precision (%)94.7
EuroEval Swedish NLU - Swerec79.56Sentiment classification Score (%)94.2
EuroEval Slovene NLU - MultiWikiQA SL70.16Reading comprehension Score (%)93.1
AbstentionBench - underspecified intent - CoCoNot/Incomprehensible - F197.92Abstention F1 (%)92.1
AbstentionBench - underspecified intent - CoCoNot/Incomprehensible - Recall95.92Abstention Recall (%)92.1
EuroEval Swedish NLU - Multi Wiki QA SV77.41Reading comprehension Score (%)89.7
EuroEval Finnish NLU - Scandisent FI92.38Sentiment classification Score (%)89.6

Interactive version: theaggregate.ai/model?slug=llama-3-1-tulu-3-8b-sft · How It Works · Data refreshed daily, snapshot 2026-09-19.