Phi-3.5-MoE-instruct — benchmark results
Microsoft's 42B mixture-of-experts Phi, activating 6.6B parameters across 16 experts, with a 128K context under MIT license (August 2024). Provider: Microsoft. Released 2024-08-23. Access: Open.
Unified ELO 1520 ± 15, rank #725 of 1776 rated models, from 50 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Stats (Social IQa) | 78 | Score (%) | 100 |
| PIQA | 88.6 | Accuracy (%) | 96.7 |
| LiveBench Cta | 60 | Score | 95.8 |
| Open CoT - LogiQA 2 | 15.14 | CoT Gain (%) | 94.7 |
| AILuminate Safety | 4 | Safety Grade (1-5) | 91.9 |
| Open CoT - LogiQA | 8.47 | CoT Gain (%) | 91.6 |
| GSM8K | 88.7 | Accuracy (%) | 89.9 |
| Open LLM Leaderboard - GPQA | 14.09 | Score | 89.6 |
| Open LLM Leaderboard - BBH | 48.77 | Score | 89.3 |
| Open LLM Leaderboard - MuSR | 17.33 | Score | 88.5 |
| Open CoT Leaderboard | 13.08 | Average CoT Gain (%) | 87 |
| Open CoT - LSAT Analytical Reasoning | 7.83 | CoT Gain (%) | 86.3 |
Interactive version: theaggregate.ai/model?slug=phi-3-5-moe-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.