Mixtral 8x22B — benchmark results
Mistral's Apache-2.0 sparse-MoE base model (141B total/39B active, 64K context) released April 2024; this row is the pretrained, non-instruct weights. Provider: Mistral. Released 2024-04-17. Access: Open.
Unified ELO 1422 ± 32, rank #1119 of 1776 rated models, from 43 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM NaturalQuestions (Closed) | 47.76 | F1 (%) | 95.6 |
| SynthPAI | 72 | Average accuracy in % | 88.2 |
| HELM NarrativeQA | 77.87 | F1 (%) | 87.2 |
| SEAL - Fortress | 56.06 | Score | 85.5 |
| HELM Lite | 75.21 | Mean win rate (self-reported) | 81.8 |
| MT-Bench PL - Extraction | 9.55 | Judge Score (0-10) | 80.6 |
| MT-Bench PL - Roleplay | 9.05 | Judge Score (0-10) | 80.6 |
| MT-Bench PL - Writing | 9.25 | Judge Score (0-10) | 80.6 |
| CyberSecEval-3 | 37.51 | Score (%) | 80 |
| HELM (Stanford) | 70.51 | Mean Win Rate (%) | 76.7 |
| MT-Bench PL - Coding | 6.45 | Judge Score (0-10) | 75.5 |
| MT-Bench PL - Math | 6.9 | Judge Score (0-10) | 75.5 |
Interactive version: theaggregate.ai/model?slug=mixtral-8x22b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.