Llama 3.3 70B Instruct [4bit]: benchmark results

Provider: Meta. Access: Open.

Unified ELO 1593 ± 25, rank #459 of 1629 rated models, from 25 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
IntentGrasp - Coronavirus Pandemic48.25Instance F1 (%) on the Coronavirus Pandemic (CP) domain of t100
IntentGrasp - E-Commerce30.01Instance F1 (%) on the E-Commerce (EC) domain of the 12,909-100
IntentGrasp - Teaching25.8Instance F1 (%) on the Teaching (T) domain of the 12,909-ite100
MTBBench-Longitudinal - Progression68.2Accuracy (%; cancer-progression questions over the evolving 100
MTBBench-Longitudinal - Recurrence56.7Accuracy (%; recurrence questions over the evolving clinical100
Open Portuguese LLM - FaQuAD NLI82.52Macro F1 (%)94.9
Open Portuguese LLM - OAB Exams58.41Accuracy (%)91.1
IntentGrasp - Policy Making34.52Instance F1 (%) on the Policy Making (PM) domain of the 12,990
IntentGrasp - Writing33.06Instance F1 (%) on the Writing (W) domain of the 12,909-item90
Open Portuguese LLM - ASSIN2 RTE93.26Macro F1 (%)87.9
Open Portuguese LLM - ENEM74.46Accuracy (%)87
MTBBench-Longitudinal66Accuracy (%; unweighted mean of the three task accuracies; 485.7

Interactive version: theaggregate.ai/model?slug=llama-3-3-70b-instruct-4bit · How It Works · Data refreshed daily, snapshot 2026-10-07.