Phi-4 Mini Instruct — benchmark results

Microsoft's MIT-licensed 3.8B dense instruct model in the Phi-4 family (February 2025), with 128K context, grouped-query attention, and a 200K-token vocabulary. Provider: Microsoft. Released 2025-02-26. Access: Open.

Unified ELO 1416 ± 9, rank #1147 of 1776 rated models, from 453 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open Arabic LLM - Aratrust Unfairness94.55Accuracy (%)91
Open LLM Leaderboard - IFEval73.78Score88.4
LatamBoard - Spanish MGSM14.4Score (%)87.9
Open Arabic LLM - Alghafa Multiple Choice Rating Sentiment NO Neutral Task80.1Accuracy (%)83.3
Open Arabic LLM - Alghafa Multiple Choice Grounded Statement Xglue Mlqa Task90.67Accuracy (%)78.7
Open LLM Leaderboard - BBH38.74Score78.5
Open Arabic LLM - Aratrust Privacy94.74Accuracy (%)78.3
Open Arabic LLM - Arabic MMLU HT Formal Logic50.79Accuracy (%)76.5
LLM Stats (Social IQa)72.5Score (%)75
SLM-RAG Arena1533.6Elo Rating75
Open LLM Leaderboard - MMLU-Pro32.58Score71.7
LatamBoard - Spanish XNLI47.47Score (%)69.7

Interactive version: theaggregate.ai/model?slug=phi-4-mini-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.