Pixtral-12B — benchmark results
Mistral's first multimodal model: a 12B Apache-2.0 vision-language model with a from-scratch 400M vision encoder and 128K context. Provider: Mistral. Released 2024-09-17. Access: Open.
Unified ELO 1452 ± 7, rank #984 of 1776 rated models, from 946 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MEGA-Bench Task - AV Human Multiview Counting | 33.3 | Task Score (%) | 100 |
| MEGA-Bench Task - MFC Bench Check Veracity | 92.9 | Task Score (%) | 100 |
| OpenVLM MMT-Bench - Relation Hallucination | 90 | Score (%) | 99.5 |
| MEGA-Bench Task - Healthcare Info Judgement | 100 | Task Score (%) | 98.8 |
| MEGA-Bench Task - Video Content Reasoning | 88.9 | Task Score (%) | 98.8 |
| OpenVLM MME - Scene | 169.5 | Score | 98.5 |
| MEGA-Bench Task - Geometry Descriptive | 35.7 | Task Score (%) | 97.7 |
| MEGA-Bench Task - Panel Images Single Question | 100 | Task Score (%) | 97.7 |
| OpenVLM MMT-Bench - HLN | 77.5 | Score (%) | 96.6 |
| MEGA-Bench Task - MFC Bench Check Clip Stable Diffusion Generate | 64.3 | Task Score (%) | 96.2 |
| OpenVLM MMT-Bench - Behavior Anomaly Detection | 70 | Score (%) | 95.6 |
| MEGA-Bench Task - Generated Video Artifacts | 43.8 | Task Score (%) | 95.3 |
Interactive version: theaggregate.ai/model?slug=pixtral-12b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.