Phi-4 Mini Instruct: benchmark results

Microsoft's MIT-licensed 3.8B dense instruct model in the Phi-4 family (February 2025), with 128K context, grouped-query attention, and a 200K-token vocabulary. Provider: Microsoft. Released 2025-02-26. Access: Open.

Unified ELO 1393 ± 1, rank #1162 of 1392 rated models, from 457 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open Arabic LLM - Aratrust Unfairness94.55Accuracy (%)91
Open LLM Leaderboard - IFEval73.78Score88.4
LatamBoard - Spanish MGSM14.4Score (%)87.9
Open Arabic LLM - Alghafa Multiple Choice Rating Sentiment NO Neutral Task80.1Accuracy (%)83.3
Open Arabic LLM - Alghafa Multiple Choice Grounded Statement Xglue Mlqa Task90.67Accuracy (%)78.7
Open LLM Leaderboard - BBH38.74Score78.5
Open Arabic LLM - Aratrust Privacy94.74Accuracy (%)78.3
Enkrypt AI - Safety Risk21.33Risk Score78.2
Open Arabic LLM - Arabic MMLU HT Formal Logic50.79Accuracy (%)76.5
LLM Stats (Social IQa)72.5Score (%)75
SLM-RAG Arena1533.6Elo Rating75
Open LLM Leaderboard - MMLU-Pro32.58Score71.7

Interactive version: theaggregate.ai/model?slug=phi-4-mini-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.