Phi-3.5-MoE-instruct — benchmark results

Microsoft's 42B mixture-of-experts Phi, activating 6.6B parameters across 16 experts, with a 128K context under MIT license (August 2024). Provider: Microsoft. Released 2024-08-23. Access: Open.

Unified ELO 1520 ± 15, rank #725 of 1776 rated models, from 50 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Stats (Social IQa)78Score (%)100
PIQA88.6Accuracy (%)96.7
LiveBench Cta60Score95.8
Open CoT - LogiQA 215.14CoT Gain (%)94.7
AILuminate Safety4Safety Grade (1-5)91.9
Open CoT - LogiQA8.47CoT Gain (%)91.6
GSM8K88.7Accuracy (%)89.9
Open LLM Leaderboard - GPQA14.09Score89.6
Open LLM Leaderboard - BBH48.77Score89.3
Open LLM Leaderboard - MuSR17.33Score88.5
Open CoT Leaderboard13.08Average CoT Gain (%)87
Open CoT - LSAT Analytical Reasoning7.83CoT Gain (%)86.3

Interactive version: theaggregate.ai/model?slug=phi-3-5-moe-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.