Phi-3 Mini 4K Instruct — benchmark results
Provider: Microsoft. Released 2024-04-23. Access: Open.
Unified ELO 1393 ± 15, rank #1259 of 1776 rated models, from 90 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HumanLikeness - Discourse-1 | 79.7 | Humanlike Score (%) | 100 |
| HumanLikeness - Syntax-1 | 89.36 | Humanlike Score (%) | 100 |
| OpenBookQA | 88 | Accuracy (%) | 98.8 |
| HumanLikeness - Discourse-2 | 76.18 | Humanlike Score (%) | 94.7 |
| HumanLikeness - Sound-2 | 67.57 | Humanlike Score (%) | 94.7 |
| HumanLikeness - Meaning-1 | 76.07 | Humanlike Score (%) | 89.5 |
| HumanLikeness - Overall | 64.73 | Overall Humanlike (%) | 89.5 |
| BiGGen-Bench | 3.82 | Average Score (1-5) | 84.3 |
| Big-Bench Hard | 71.7 | Average (%) | 82.3 |
| ARC Challenge (AI2) | 84.9 | Accuracy (%) | 82.1 |
| Open LLM Leaderboard - GPQA | 10.96 | Score | 81.4 |
| Open CoT - LSAT Analytical Reasoning | 6.52 | CoT Gain (%) | 81.3 |
Interactive version: theaggregate.ai/model?slug=phi-3-mini-4k-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.