Llama 3.2 90B Instruct — benchmark results
Meta's largest open vision model: Llama 3.1 70B with a cross-attention image adapter, tuned for visual reasoning and captioning (September 2024). Provider: Meta. Released 2024-09-25. Access: Open.
Unified ELO 1504 ± 22, rank #784 of 1776 rated models, from 26 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Stats (AI2D) | 92.3 | Score (%) | 71 |
| BALROG BabaIsAI (LLM) | 43.9 | Progress (%) | 69.7 |
| BALROG BabyAI (VLM) | 66 | Progress (%) | 63.6 |
| BALROG BabyAI (LLM) | 72 | Progress (%) | 62.1 |
| BALROG BabaIsAI (VLM) | 21.9 | Progress (%) | 60 |
| BALROG Crafter (LLM) | 31.7 | Progress (%) | 57.6 |
| LLM Stats (ChartQA) | 85.5 | Score (%) | 47.8 |
| BALROG TextWorld (LLM) | 11.2 | Progress (%) | 37.9 |
| BALROG MiniHack (LLM) | 5 | Progress (%) | 31.8 |
| AA Humanity's Last Exam | 4.92 | Accuracy (%) | 31.6 |
| AA MMLU-Pro | 67.1 | Accuracy (%) | 29.7 |
| LLM Stats (TextVQA) | 73.5 | Score (%) | 28.6 |
Interactive version: theaggregate.ai/model?slug=llama-3-2-90b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.