Phi-3 Mini Instruct 3.8B — benchmark results
Provider: Microsoft. Released 2024-04-23. Access: Open.
Unified ELO 1113 ± 114, rank #1821 of 1841 rated models, from 13 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA Humanity's Last Exam | 4.43 | Accuracy (%) | 20.9 |
| AA-LCR | 2 | Score (self-reported) | 19.1 |
| AA Long Context Reasoning | 2 | Accuracy (%) | 14.2 |
| Artificial Analysis Intelligence Index | 4.59 | Intelligence Index | 13.5 |
| AA MATH-500 | 45.67 | Accuracy (%) | 12.1 |
| AA MMLU-Pro | 43.51 | Accuracy (%) | 10.8 |
| AA LiveCodeBench | 11.64 | Pass@1 (%) | 10.5 |
| AA GPQA Diamond | 31.92 | Accuracy (%) | 9.2 |
| Epoch AI - Scicode | 9.03 | Score | 8.6 |
| AA IFBench | 23.88 | Accuracy (%) | 6.2 |
| AA Terminal-Bench Hard | 0 | Accuracy (%) | 5.5 |
| AA TAU-2 Bench | 0 | Accuracy (%) | 2.5 |
Interactive version: theaggregate.ai/model?slug=phi-3-mini-instruct-3-8b · How It Works · Data refreshed daily, snapshot 2026-07-25.