WizardLM-2 8x22B — benchmark results

Microsoft's Mixtral-8x22B finetune briefly pulled days after its April 2024 release over missing toxicity testing, surviving via community mirrors. Provider: Microsoft. Released 2024-04-15. Access: Open.

Unified ELO 1511 ± 18, rank #764 of 1776 rated models, from 44 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard - GPQA17.56Score95.3
LogicKor - Reasoning8.5Score (0-10)90.5
LogicKor - Single-Turn8.52Score (0-10)89.6
Open Chinese LLM - GSM8K64.82Accuracy (%)89
Open LLM Leaderboard - BBH48.58Score88.9
LogicKor - Writing9.35Score (0-10)87.8
LogicKor8.13Score (0-10)85.8
LogicKor - Coding9.28Score (0-10)85.8
Open LLM Leaderboard - MMLU-Pro39.96Score84.9
LogicKor - Multi-Turn7.73Score (0-10)82.5
Open PL LLM - Generative66.42Average Generative Score (%)82.2
Open PL LLM Leaderboard62.35Average Score (%)80.4

Interactive version: theaggregate.ai/model?slug=wizardlm-2-8x22b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.