Phi-4 Mini Instruct — benchmark results
Microsoft's MIT-licensed 3.8B dense instruct model in the Phi-4 family (February 2025), with 128K context, grouped-query attention, and a 200K-token vocabulary. Provider: Microsoft. Released 2025-02-26. Access: Open.
Unified ELO 1416 ± 9, rank #1147 of 1776 rated models, from 453 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open Arabic LLM - Aratrust Unfairness | 94.55 | Accuracy (%) | 91 |
| Open LLM Leaderboard - IFEval | 73.78 | Score | 88.4 |
| LatamBoard - Spanish MGSM | 14.4 | Score (%) | 87.9 |
| Open Arabic LLM - Alghafa Multiple Choice Rating Sentiment NO Neutral Task | 80.1 | Accuracy (%) | 83.3 |
| Open Arabic LLM - Alghafa Multiple Choice Grounded Statement Xglue Mlqa Task | 90.67 | Accuracy (%) | 78.7 |
| Open LLM Leaderboard - BBH | 38.74 | Score | 78.5 |
| Open Arabic LLM - Aratrust Privacy | 94.74 | Accuracy (%) | 78.3 |
| Open Arabic LLM - Arabic MMLU HT Formal Logic | 50.79 | Accuracy (%) | 76.5 |
| LLM Stats (Social IQa) | 72.5 | Score (%) | 75 |
| SLM-RAG Arena | 1533.6 | Elo Rating | 75 |
| Open LLM Leaderboard - MMLU-Pro | 32.58 | Score | 71.7 |
| LatamBoard - Spanish XNLI | 47.47 | Score (%) | 69.7 |
Interactive version: theaggregate.ai/model?slug=phi-4-mini-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.