PhiBench: leaderboard

Internal benchmark of language-model skills and reasoning, covering coding (debugging, extending incomplete code, explaining code snippets) and mathematics (identifying proof errors, generating related problems). Created by Microsoft's research team to address limitations of standard academic benchmarks and guide the development of the Phi-4 model.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturated. 7 models tracked.

Top models

#ModelScore
1GPT-4o72.4
2Qwen 2.5 72B Instruct64.6
3GPT-4o Mini58.7
4Llama 3.3 70B Instruct57.1
5phi-4-14B56.2
6Qwen 2.5 14B Instruct49.8

Interactive version: theaggregate.ai/benchmark?slug=phibench · How It Works · Data refreshed daily, snapshot 2026-09-05.