Phi-1: benchmark results

Provider: Microsoft. Released 2023-06-20. Access: Open.

Unified ELO 1261 ± 1, rank #1379 of 1392 rated models, from 19 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AI Energy Score (Text Generation)5Energy Score (1-5)79.5
BigCode Models Leaderboard51.2HumanEval Python Pass@1 (%)62.7
Open LLM Leaderboard - GPQA2.01Score19.3
Open LLM Leaderboard - MuSR3.7Score17.7
Open LLM Leaderboard - IFEval20.68Score13
RABBITS B4BQA25.76Accuracy (%)10.7
Open LLM Leaderboard - BBH4.27Score8.3
LLMZSZL Leaderboard25.73Score8.2
Open LLM Leaderboard - MMLU-Pro1.8Score7.4
ShaderMatch1.52Clone Match Rate (%)7.1
Open LLM Leaderboard - MATH Level 50.98Score7
InfiBench14.28Score (%)2.9

Interactive version: theaggregate.ai/model?slug=phi-1 · How It Works · Data refreshed daily, snapshot 2026-09-05.