Phi-1 — benchmark results
Provider: Microsoft. Released 2023-06-20. Access: Open.
Unified ELO 1080 ± 49, rank #1767 of 1776 rated models, from 20 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AI Energy Score (Text Generation) | 5 | Energy Score (1-5) | 79.5 |
| BigCode Models Leaderboard | 51.2 | HumanEval Python Pass@1 (%) | 62.7 |
| Open LLM Leaderboard - GPQA | 2.01 | Score | 19.3 |
| Open LLM Leaderboard - MuSR | 3.7 | Score | 17.7 |
| Open LLM Leaderboard - IFEval | 20.68 | Score | 13 |
| RABBITS B4BQA | 25.76 | Accuracy (%) | 12 |
| Open LLM Leaderboard - BBH | 4.27 | Score | 8.3 |
| LLMZSZL Leaderboard | 25.73 | Score | 8.2 |
| Open LLM Leaderboard - MMLU-Pro | 1.8 | Score | 7.4 |
| ShaderMatch | 1.52 | Clone Match Rate (%) | 7.1 |
| Open LLM Leaderboard - MATH Level 5 | 0.98 | Score | 7 |
| InfiBench | 14.28 | Score (%) | 2.9 |
Interactive version: theaggregate.ai/model?slug=phi-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.