Phi-3 Mini 128K Instruct — benchmark results
Provider: Microsoft. Released 2024-04-23. Access: Open.
Unified ELO 1386 ± 15, rank #1281 of 1776 rated models, from 66 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - BBH | 37.1 | Score | 76.6 |
| BiGGen-Bench | 3.68 | Average Score (1-5) | 75.5 |
| Open LLM Leaderboard - GPQA | 9.06 | Score | 73.8 |
| Open LLM Leaderboard - IFEval | 59.76 | Score | 71.7 |
| LiveBench Cta | 52 | Score | 65.3 |
| Open LLM Leaderboard - MMLU-Pro | 30.38 | Score | 61.8 |
| Open LLM Leaderboard - MATH Level 5 | 14.05 | Score | 60.3 |
| EuroEval Portuguese NLU - ScaLA PT | 16.29 | Linguistic acceptability Score (%) | 58.4 |
| LiveBench Table Reformat | 44 | Score | 54.2 |
| Open Chinese LLM - TruthfulQA MC | 53.96 | Accuracy (%) | 53.6 |
| SpeechMap Compliance | 64.3 | % Requests Completed | 52.7 |
| BABILong (NIAH) | 45.4 | Avg Accuracy (%) | 51.7 |
Interactive version: theaggregate.ai/model?slug=phi-3-mini-128k-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.