Pixtral-12B — benchmark results

Mistral's first multimodal model: a 12B Apache-2.0 vision-language model with a from-scratch 400M vision encoder and 128K context. Provider: Mistral. Released 2024-09-17. Access: Open.

Unified ELO 1452 ± 7, rank #984 of 1776 rated models, from 946 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
MEGA-Bench Task - AV Human Multiview Counting33.3Task Score (%)100
MEGA-Bench Task - MFC Bench Check Veracity92.9Task Score (%)100
OpenVLM MMT-Bench - Relation Hallucination90Score (%)99.5
MEGA-Bench Task - Healthcare Info Judgement100Task Score (%)98.8
MEGA-Bench Task - Video Content Reasoning88.9Task Score (%)98.8
OpenVLM MME - Scene169.5Score98.5
MEGA-Bench Task - Geometry Descriptive35.7Task Score (%)97.7
MEGA-Bench Task - Panel Images Single Question100Task Score (%)97.7
OpenVLM MMT-Bench - HLN77.5Score (%)96.6
MEGA-Bench Task - MFC Bench Check Clip Stable Diffusion Generate64.3Task Score (%)96.2
OpenVLM MMT-Bench - Behavior Anomaly Detection70Score (%)95.6
MEGA-Bench Task - Generated Video Artifacts43.8Task Score (%)95.3

Interactive version: theaggregate.ai/model?slug=pixtral-12b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.