Mixtral 8x7B — benchmark results
Mistral's first Apache-2.0 sparse-MoE model (46.7B total/12.9B active, 8 experts pick 2), December 2023; this row is the pretrained base, non-instruct weights. Provider: Mistral. Released 2023-12-11. Access: Open.
Unified ELO 1420 ± 26, rank #1128 of 1776 rated models, from 35 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| XDomainBench | 56.4 | R (k=1, Deterministic) (self-reported) | 92.3 |
| MT-Bench PL - Writing | 9.35 | Judge Score (0-10) | 90.8 |
| MT-Bench PL - Roleplay | 8.95 | Judge Score (0-10) | 75.5 |
| AgentClinic | 37.1 | AgentClinic-MedQA diagnostic accuracy (self-reported) | 66.7 |
| MT-Bench PL - Humanities | 9.45 | Judge Score (0-10) | 61.2 |
| MT-Bench PL - Overall | 7.64 | Judge Score (0-10) | 57.1 |
| MT-Bench PL - Coding | 5.2 | Judge Score (0-10) | 55.1 |
| MT-Bench PL - Math | 5.65 | Judge Score (0-10) | 55.1 |
| MT-Bench PL - Reasoning | 5.8 | Judge Score (0-10) | 55.1 |
| MT-Bench PL - STEM | 8.55 | Judge Score (0-10) | 49 |
| MT-Bench PL - Extraction | 8.15 | Judge Score (0-10) | 48 |
| CRUXEval | 40.5 | Output Prediction pass@1 (%) | 42.9 |
Interactive version: theaggregate.ai/model?slug=mixtral-8x7b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.