Phi-3-medium-4k-instruct — benchmark results
Microsoft's 14B Phi-3 Medium instruct model with a 4K context, trained on synthetic and filtered web data. Provider: Microsoft. Released 2024-05-21. Access: Open.
Unified ELO 1444 ± 13, rank #1028 of 1776 rated models, from 96 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LatamBoard - Spanish Escola | 74.55 | Score (%) | 100 |
| LatamBoard - Spanish OpenBookQA | 40 | Score (%) | 100 |
| LatamBoard - Spanish WNLI | 81.69 | Score (%) | 100 |
| Shlepa - Books MC (RU) | 38.17 | Accuracy (%) | 100 |
| LatamBoard - Spanish MGSM | 27.6 | Score (%) | 97 |
| LatamBoard - Spanish Score | 64.9 | Score (%) | 97 |
| LatamBoard - Spanish COPA | 84.6 | Score (%) | 93.9 |
| Pinocchio Italian - Cultura | 71.25 | Accuracy (%) | 93.2 |
| Pinocchio Italian - Diritto | 61.53 | Accuracy (%) | 93.2 |
| Pinocchio Italian - Generale | 67.38 | Accuracy (%) | 93.2 |
| Pinocchio Italian - Logica | 53.26 | Accuracy (%) | 93.2 |
| Pinocchio Italian - Matematica E Scienze | 64.17 | Accuracy (%) | 93.2 |
Interactive version: theaggregate.ai/model?slug=phi-3-medium-4k-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.