Mixtral 8x7B — benchmark results

Mistral's first Apache-2.0 sparse-MoE model (46.7B total/12.9B active, 8 experts pick 2), December 2023; this row is the pretrained base, non-instruct weights. Provider: Mistral. Released 2023-12-11. Access: Open.

Unified ELO 1420 ± 26, rank #1128 of 1776 rated models, from 35 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
XDomainBench56.4R (k=1, Deterministic) (self-reported)92.3
MT-Bench PL - Writing9.35Judge Score (0-10)90.8
MT-Bench PL - Roleplay8.95Judge Score (0-10)75.5
AgentClinic37.1AgentClinic-MedQA diagnostic accuracy (self-reported)66.7
MT-Bench PL - Humanities9.45Judge Score (0-10)61.2
MT-Bench PL - Overall7.64Judge Score (0-10)57.1
MT-Bench PL - Coding5.2Judge Score (0-10)55.1
MT-Bench PL - Math5.65Judge Score (0-10)55.1
MT-Bench PL - Reasoning5.8Judge Score (0-10)55.1
MT-Bench PL - STEM8.55Judge Score (0-10)49
MT-Bench PL - Extraction8.15Judge Score (0-10)48
CRUXEval40.5Output Prediction pass@1 (%)42.9

Interactive version: theaggregate.ai/model?slug=mixtral-8x7b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.