Phi-3-medium-128k-instruct — benchmark results
Microsoft's 14B Phi-3 Medium instruct model with a 128K context, trained on synthetic and filtered web data. Provider: Microsoft. Released 2024-05-21. Access: Open.
Unified ELO 1451 ± 16, rank #993 of 1776 rated models, from 47 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| OpenBookQA | 87.4 | Accuracy (%) | 95.2 |
| Big-Bench Hard | 81.4 | Average (%) | 93.8 |
| ARC Challenge (AI2) | 91.6 | Accuracy (%) | 93.6 |
| Open Korean LLM Leaderboard | 568.43 | Average Score (%) | 92.4 |
| Open LLM Leaderboard - BBH | 48.46 | Score | 88.7 |
| Open LLM Leaderboard - MMLU-Pro | 41.24 | Score | 86 |
| CanAiCode | 100 | Junior-v2 Python Pass Rate (%) | 85.6 |
| Open LLM Leaderboard - GPQA | 11.52 | Score | 83.1 |
| WinoGrande | 81.5 | Accuracy (%) | 79.4 |
| MMLU | 78 | Accuracy (%) | 76.3 |
| BABILong (NIAH) | 57.3 | Avg Accuracy (%) | 75.9 |
| HellaSwag | 82.4 | Accuracy (%) | 72.4 |
Interactive version: theaggregate.ai/model?slug=phi-3-medium-128k-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.