Mixtral 8x22B — benchmark results

Mistral's Apache-2.0 sparse-MoE base model (141B total/39B active, 64K context) released April 2024; this row is the pretrained, non-instruct weights. Provider: Mistral. Released 2024-04-17. Access: Open.

Unified ELO 1422 ± 32, rank #1119 of 1776 rated models, from 43 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM NaturalQuestions (Closed)47.76F1 (%)95.6
SynthPAI72Average accuracy in %88.2
HELM NarrativeQA77.87F1 (%)87.2
SEAL - Fortress56.06Score85.5
HELM Lite75.21Mean win rate (self-reported)81.8
MT-Bench PL - Extraction9.55Judge Score (0-10)80.6
MT-Bench PL - Roleplay9.05Judge Score (0-10)80.6
MT-Bench PL - Writing9.25Judge Score (0-10)80.6
CyberSecEval-337.51Score (%)80
HELM (Stanford)70.51Mean Win Rate (%)76.7
MT-Bench PL - Coding6.45Judge Score (0-10)75.5
MT-Bench PL - Math6.9Judge Score (0-10)75.5

Interactive version: theaggregate.ai/model?slug=mixtral-8x22b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.