Llama 3.2 90B Instruct: benchmark results
Meta's largest open vision model: Llama 3.1 70B with a cross-attention image adapter, tuned for visual reasoning and captioning (September 2024). Provider: Meta. Released 2024-09-25. Access: Open.
Unified ELO 1477 ± 1, rank #801 of 1392 rated models, from 26 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BALROG BabyAI (VLM) | 66 | Progress (%) | 77.8 |
| LLM Stats (AI2D) | 92.3 | Score (%) | 72.7 |
| BALROG BabaIsAI (LLM) | 43.9 | Progress (%) | 69.7 |
| BALROG BabaIsAI (VLM) | 21.9 | Progress (%) | 66.7 |
| BALROG BabyAI (LLM) | 72 | Progress (%) | 62.1 |
| BALROG Crafter (LLM) | 31.7 | Progress (%) | 57.6 |
| LLM Stats (ChartQA) | 85.5 | Score (%) | 52 |
| BALROG TextWorld (LLM) | 11.2 | Progress (%) | 37.9 |
| BALROG MiniHack (LLM) | 5 | Progress (%) | 31.8 |
| BALROG NetHack (VLM) | 0 | Progress (%) | 27.8 |
| LLM Stats (MathVista) | 57.3 | Score (%) | 27 |
| LLM Stats (TextVQA) | 73.5 | Score (%) | 26.7 |
Interactive version: theaggregate.ai/model?slug=llama-3-2-90b-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.