Mixtral 8x7B: benchmark results

Mistral's first Apache-2.0 sparse-MoE model (46.7B total/12.9B active, 8 experts pick 2), December 2023; this row is the pretrained base, non-instruct weights. Provider: Mistral. Released 2023-12-11. Access: Open.

Unified ELO 1452 ± 1, rank #935 of 1392 rated models, from 35 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
XDomainBench56.4R (k=1, Deterministic) (self-reported)92.3
MT-Bench PL - Writing9.35Judge Score (0-10)90.8
MT-Bench PL - Roleplay8.95Judge Score (0-10)75.5
AgentClinic37.1AgentClinic-MedQA diagnostic accuracy (self-reported)66.7
MT-Bench PL - Humanities9.45Judge Score (0-10)61.2
MT-Bench PL - Overall7.64Judge Score (0-10)57.1
AutoRace27Average Score (%, six reasoning tasks)55.6
MT-Bench PL - Coding5.2Judge Score (0-10)55.1
MT-Bench PL - Math5.65Judge Score (0-10)55.1
MT-Bench PL - Reasoning5.8Judge Score (0-10)55.1
MT-Bench PL - STEM8.55Judge Score (0-10)49
MT-Bench PL - Extraction8.15Judge Score (0-10)48

Interactive version: theaggregate.ai/model?slug=mixtral-8x7b · How It Works · Data refreshed daily, snapshot 2026-09-05.