Llama 3.3 70B — benchmark results

Meta Llama 3.3 70B model row. Provider: Meta. Released 2024-12-06. Access: Open.

Unified ELO 1516 ± 18, rank #745 of 1776 rated models, from 47 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
NoLiMa97.3Base Score (%)95.2
TextClass Benchmark1746.41Meta-Elo (self-reported)92.5
ASCIIEval32.74Macro-Avg Accuracy (%)80.6
Divergent Thinking4.44Divergence Score77.8
SEAL - Fortress44.79Score69.1
SEA-HELM60.42Mean Score (%)58.1
Confabulation Leaderboard (Lechmazur)17.82Confabulation rate % (lower is better)57.9
AI for Education Pedagogy - Technology79.25Accuracy (%)55.2
Halluverse-M370.53Macro Accuracy (self-reported)53.8
MMTU - Data Cleaning38.64Accuracy (%)52
MMTU - KB mapping43.55Accuracy (%)52
gg-bench17.77Games Won50

Interactive version: theaggregate.ai/model?slug=llama-3-3-70b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.