Llama 3.1 Tulu3 405B — benchmark results
Provider: Meta. Released 2025-01-30. Access: Open.
Unified ELO 1486 ± 38, rank #902 of 1841 rated models, from 7 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA MATH-500 | 77.8 | Accuracy (%) | 43.9 |
| Epoch AI - Scicode | 30.21 | Score | 42.4 |
| AA MMLU-Pro | 71.62 | Accuracy (%) | 39.5 |
| AA LiveCodeBench | 29.1 | Pass@1 (%) | 33.5 |
| AA GPQA Diamond | 51.62 | Accuracy (%) | 28.3 |
| Artificial Analysis Intelligence Index | 8.28 | Intelligence Index | 25.9 |
| AA Humanity's Last Exam | 3.46 | Accuracy (%) | 3.8 |
Interactive version: theaggregate.ai/model?slug=llama-3-1-tulu3-405b · How It Works · Data refreshed daily, snapshot 2026-07-25.