Llama 3.3 70B IT — benchmark results
Meta's text-only 70B instruct tune that matches Llama 3.1 405B on key benchmarks at a fraction of the cost, with a 128K context (December 2024). Provider: Meta. Released 2024-12-06. Access: Open.
Unified ELO 1514 ± 20, rank #756 of 1776 rated models, from 7 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BALROG BabyAI (LLM) | 66 | Progress (%) | 50 |
| BALROG Crafter (LLM) | 28.6 | Progress (%) | 48.5 |
| BALROG NetHack (LLM) | 0.4 | Progress (%) | 48.5 |
| BALROG BabaIsAI (LLM) | 29.2 | Progress (%) | 42.4 |
| BALROG MiniHack (LLM) | 5 | Progress (%) | 31.8 |
| BALROG TextWorld (LLM) | 9 | Progress (%) | 30.3 |
| LEXam | 28.19 | Multiple-Choice Accuracy (%) | 13.3 |
Interactive version: theaggregate.ai/model?slug=llama-3-3-70b-it · How the rankings work · Data refreshed daily, snapshot 2026-07-22.