Llama 3.3 70B — benchmark results
Meta Llama 3.3 70B model row. Provider: Meta. Released 2024-12-06. Access: Open.
Unified ELO 1516 ± 18, rank #745 of 1776 rated models, from 47 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| NoLiMa | 97.3 | Base Score (%) | 95.2 |
| TextClass Benchmark | 1746.41 | Meta-Elo (self-reported) | 92.5 |
| ASCIIEval | 32.74 | Macro-Avg Accuracy (%) | 80.6 |
| Divergent Thinking | 4.44 | Divergence Score | 77.8 |
| SEAL - Fortress | 44.79 | Score | 69.1 |
| SEA-HELM | 60.42 | Mean Score (%) | 58.1 |
| Confabulation Leaderboard (Lechmazur) | 17.82 | Confabulation rate % (lower is better) | 57.9 |
| AI for Education Pedagogy - Technology | 79.25 | Accuracy (%) | 55.2 |
| Halluverse-M3 | 70.53 | Macro Accuracy (self-reported) | 53.8 |
| MMTU - Data Cleaning | 38.64 | Accuracy (%) | 52 |
| MMTU - KB mapping | 43.55 | Accuracy (%) | 52 |
| gg-bench | 17.77 | Games Won | 50 |
Interactive version: theaggregate.ai/model?slug=llama-3-3-70b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.