Llama 3.3 70B Instruct — benchmark results

Meta Llama 3.3 70B instruction-tuned checkpoint. Provider: Meta. Released 2024-12-06. Access: Open.

Unified ELO 1543 ± 5, rank #642 of 1776 rated models, from 1110 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BlueBench - Product Help86.64Score (%)100
BlueBench - Summarization19.47Score (%)100
DiagFlowBench85Step Accuracy (self-reported)100
EVALITA - evalita NER44.58CPS100
EVALITA - hate-speech-detection77.82CPS100
EuroEval Estonian Summarization - ERR News32.83Score (%)100
EuroEval Norwegian Summarization - NO Sammendrag32.63Score (%)100
French LLM Leaderboard - IFEval FR74.51Score (%)100
HELM SeaHELM - LINDSEA Scalar Implicatures (id)92.31EM100
Medical Information Response Audit (MIRA)2.95Underinformative simplification (D3) (self-reported)100
Open LLM Leaderboard - IFEval89.98Score100
Open Persian LLM - AUT Multiple Choice Persian71.4Accuracy (%)100

Interactive version: theaggregate.ai/model?slug=llama-3-3-70b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.