Phi-1 — benchmark results

Provider: Microsoft. Released 2023-06-20. Access: Open.

Unified ELO 1080 ± 49, rank #1767 of 1776 rated models, from 20 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AI Energy Score (Text Generation)5Energy Score (1-5)79.5
BigCode Models Leaderboard51.2HumanEval Python Pass@1 (%)62.7
Open LLM Leaderboard - GPQA2.01Score19.3
Open LLM Leaderboard - MuSR3.7Score17.7
Open LLM Leaderboard - IFEval20.68Score13
RABBITS B4BQA25.76Accuracy (%)12
Open LLM Leaderboard - BBH4.27Score8.3
LLMZSZL Leaderboard25.73Score8.2
Open LLM Leaderboard - MMLU-Pro1.8Score7.4
ShaderMatch1.52Clone Match Rate (%)7.1
Open LLM Leaderboard - MATH Level 50.98Score7
InfiBench14.28Score (%)2.9

Interactive version: theaggregate.ai/model?slug=phi-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.