Phi-3.5-mini-instruct — benchmark results
Microsoft's 3.8B MIT-licensed small model with a 128K context and improved multilingual coverage. Provider: Microsoft. Released 2024-08-20. Access: Open.
Unified ELO 1436 ± 12, rank #1061 of 1776 rated models, from 146 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Stats (SQuALITY) | 24.3 | Score (%) | 100 |
| LIBRA - MatreshkaYesNo | 73.42 | Dataset Total Score (%) | 91.7 |
| Open Japanese LLM - Xlsum JA Bert Score JA F1 | 70.67 | Score (%) | 88.8 |
| LLM Stats (Social IQa) | 74.7 | Score (%) | 87.5 |
| GSM8K | 86.2 | Accuracy (%) | 85.1 |
| Open LLM Leaderboard - GPQA | 11.97 | Score | 84.3 |
| Open Japanese LLM - Xlsum JA Rouge1 | 29.95 | Score (%) | 80.3 |
| Open Japanese LLM - Xlsum JA RougeLsum | 25.34 | Score (%) | 79.8 |
| Open Japanese LLM - SUM | 10.73 | Score (%) | 76.6 |
| Open LLM Leaderboard - BBH | 36.75 | Score | 76 |
| LIBRA - ruBABILongQA5 | 71.22 | Dataset Total Score (%) | 75 |
| LIBRA - ruQuALITY | 81.3 | Dataset Total Score (%) | 75 |
Interactive version: theaggregate.ai/model?slug=phi-3-5-mini-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.