Llama 3.2 11B Instruct — benchmark results
Meta Llama 3.2 11B vision instruction-tuned checkpoint. Provider: Meta. Released 2024-09-25. Access: Open.
Unified ELO 1394 ± 7, rank #1253 of 1776 rated models, from 639 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - schnovel | 1.91 | Dataset z-score | 98.8 |
| Open LMM Reasoning - WeMath - InsufficientKnowledge (Loose) | 80 | Accuracy (%) | 97.5 |
| Open LMM Reasoning - WeMath - RoteMemorization (Loose) | 36 | Accuracy (%) | 97.5 |
| OpenVLM MMT-Bench - Abstract Visual Recognition | 90 | Score (%) | 96.8 |
| AGC-Bench - puntuguese | 1.02 | Dataset z-score | 96.3 |
| OpenVLM MMT-Bench - National Flag Recognition | 100 | Score (%) | 94.9 |
| OpenVLM MMT-Bench - Table Structure Recognition | 100 | Score (%) | 94.7 |
| OpenVLM MMT-Bench - Logo and Brand Recognition | 100 | Score (%) | 94.4 |
| AGC-Bench - creatset | 1.47 | Dataset z-score | 93.9 |
| OpenVLM MMT-Bench - Humanities and Social Science | 72.7 | Score (%) | 93.9 |
| OpenVLM MMT-Bench - Astronomical Recognition | 88.9 | Score (%) | 93 |
| AGC-Bench - slang_generation | 1.1 | Dataset z-score | 92.7 |
Interactive version: theaggregate.ai/model?slug=llama-3-2-11b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.