Phi-3-medium-4k-instruct — benchmark results

Microsoft's 14B Phi-3 Medium instruct model with a 4K context, trained on synthetic and filtered web data. Provider: Microsoft. Released 2024-05-21. Access: Open.

Unified ELO 1444 ± 13, rank #1028 of 1776 rated models, from 96 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LatamBoard - Spanish Escola74.55Score (%)100
LatamBoard - Spanish OpenBookQA40Score (%)100
LatamBoard - Spanish WNLI81.69Score (%)100
Shlepa - Books MC (RU)38.17Accuracy (%)100
LatamBoard - Spanish MGSM27.6Score (%)97
LatamBoard - Spanish Score64.9Score (%)97
LatamBoard - Spanish COPA84.6Score (%)93.9
Pinocchio Italian - Cultura71.25Accuracy (%)93.2
Pinocchio Italian - Diritto61.53Accuracy (%)93.2
Pinocchio Italian - Generale67.38Accuracy (%)93.2
Pinocchio Italian - Logica53.26Accuracy (%)93.2
Pinocchio Italian - Matematica E Scienze64.17Accuracy (%)93.2

Interactive version: theaggregate.ai/model?slug=phi-3-medium-4k-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.