Phi-3-medium-4k-instruct: benchmark results

Microsoft's 14B Phi-3 Medium instruct model with a 4K context, trained on synthetic and filtered web data. Provider: Microsoft. Released 2024-05-21. Access: Open.

Unified ELO 1477 ± 1, rank #797 of 1392 rated models, from 122 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LatamBoard - Spanish Escola74.55Score (%)100
LatamBoard - Spanish OpenBookQA40Score (%)100
LatamBoard - Spanish WNLI81.69Score (%)100
Shlepa - Books MC (RU)38.17Accuracy (%)100
LatamBoard - Spanish MGSM27.6Score (%)97
LatamBoard - Spanish Score64.9Score (%)97
LatamBoard - Spanish COPA84.6Score (%)93.9
Pinocchio Italian - Cultura71.25Accuracy (%)93.2
Pinocchio Italian - Diritto61.53Accuracy (%)93.2
Pinocchio Italian - Generale67.38Accuracy (%)93.2
Pinocchio Italian - Logica53.26Accuracy (%)93.2
Pinocchio Italian - Matematica E Scienze64.17Accuracy (%)93.2

Interactive version: theaggregate.ai/model?slug=phi-3-medium-4k-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.