Llama 3.3 70B Instruct [4bit]: benchmark results
Provider: Meta. Access: Open.
Unified ELO 1593 ± 25, rank #459 of 1629 rated models, from 25 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| IntentGrasp - Coronavirus Pandemic | 48.25 | Instance F1 (%) on the Coronavirus Pandemic (CP) domain of t | 100 |
| IntentGrasp - E-Commerce | 30.01 | Instance F1 (%) on the E-Commerce (EC) domain of the 12,909- | 100 |
| IntentGrasp - Teaching | 25.8 | Instance F1 (%) on the Teaching (T) domain of the 12,909-ite | 100 |
| MTBBench-Longitudinal - Progression | 68.2 | Accuracy (%; cancer-progression questions over the evolving | 100 |
| MTBBench-Longitudinal - Recurrence | 56.7 | Accuracy (%; recurrence questions over the evolving clinical | 100 |
| Open Portuguese LLM - FaQuAD NLI | 82.52 | Macro F1 (%) | 94.9 |
| Open Portuguese LLM - OAB Exams | 58.41 | Accuracy (%) | 91.1 |
| IntentGrasp - Policy Making | 34.52 | Instance F1 (%) on the Policy Making (PM) domain of the 12,9 | 90 |
| IntentGrasp - Writing | 33.06 | Instance F1 (%) on the Writing (W) domain of the 12,909-item | 90 |
| Open Portuguese LLM - ASSIN2 RTE | 93.26 | Macro F1 (%) | 87.9 |
| Open Portuguese LLM - ENEM | 74.46 | Accuracy (%) | 87 |
| MTBBench-Longitudinal | 66 | Accuracy (%; unweighted mean of the three task accuracies; 4 | 85.7 |
Interactive version: theaggregate.ai/model?slug=llama-3-3-70b-instruct-4bit · How It Works · Data refreshed daily, snapshot 2026-10-07.