Phi-4-mini (Reasoning) — benchmark results

Provider: Microsoft. Released 2025-04-30. Access: Open.

Unified ELO 1237 ± 21, rank #1646 of 1776 rated models, from 109 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HarmActionsEval2.84SafeActions@1 (self-reported)72.2
EuroEval English NLU - CoNLL EN69.04Named entity recognition Score (%)50.6
EuroEval Portuguese Knowledge39.64Knowledge Average Score (%)48
EuroEval French Knowledge41.54Knowledge Average Score (%)43.7
ZeroEval MATH-50094.6MATH-500 Score41.9
EuroEval German Knowledge38.36Knowledge Average Score (%)41.8
EuroEval Spanish Knowledge35.22Knowledge Average Score (%)41.1
EuroEval Swedish Knowledge30.63Knowledge Average Score (%)40.4
EuroEval Spanish NLU - ScaLA ES5.8Linguistic acceptability Score (%)39.9
EuroEval Dutch Knowledge35.12Knowledge Average Score (%)39.2
EuroEval Italian Knowledge32.58Knowledge Average Score (%)39.1
EuroEval Icelandic NLU - ScaLA IS1.17Linguistic acceptability Score (%)37.7

Interactive version: theaggregate.ai/model?slug=phi-4-mini-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.