Phi-2 — benchmark results
Provider: Microsoft. Released 2023-12-12. Access: Open.
Unified ELO 1251 ± 17, rank #1612 of 1776 rated models, from 111 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open CoT - LogiQA | 6.23 | CoT Gain (%) | 79.4 |
| OpenBookQA | 73.6 | Accuracy (%) | 78.6 |
| Open LLM Leaderboard - MuSR | 13.84 | Score | 76.2 |
| Big-Bench Hard | 59.4 | Average (%) | 72.9 |
| NPHardEval - GCP D | 55.82 | Accuracy (%) | 72.7 |
| ARC Challenge (AI2) | 75.9 | Accuracy (%) | 71.8 |
| URIAL-Bench - Coding | 4.25 | Judge Score (0-10) | 66.7 |
| URIAL-Bench - Humanities | 8.85 | Judge Score (0-10) | 66.7 |
| URIAL-Bench - Math | 3.8 | Judge Score (0-10) | 66.7 |
| Open CoT - LogiQA 2 | 8.65 | CoT Gain (%) | 60.3 |
| EvoEval Combine | 14 | Pass@1 (%) | 58 |
| Open CoT Leaderboard | 8.08 | Average CoT Gain (%) | 56.5 |
Interactive version: theaggregate.ai/model?slug=phi-2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.