PhiBench: leaderboard
Internal benchmark of language-model skills and reasoning, covering coding (debugging, extending incomplete code, explaining code snippets) and mathematics (identifying proof errors, generating related problems). Created by Microsoft's research team to address limitations of standard academic benchmarks and guide the development of the Phi-4 model.
Metric: Score (self-reported). Source: benchmarklist.com. Status: saturated. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o | 72.4 |
| 2 | Qwen 2.5 72B Instruct | 64.6 |
| 3 | GPT-4o Mini | 58.7 |
| 4 | Llama 3.3 70B Instruct | 57.1 |
| 5 | phi-4-14B | 56.2 |
| 6 | Qwen 2.5 14B Instruct | 49.8 |
Interactive version: theaggregate.ai/benchmark?slug=phibench · How It Works · Data refreshed daily, snapshot 2026-09-05.