phi-4-14B — benchmark results
Microsoft's 14B Phi-4 trained heavily on synthetic data, with weights released on Hugging Face under MIT in January 2025. Provider: Microsoft. Released 2025-01-16. Access: Open.
Unified ELO 1425 ± 93, rank #1107 of 1776 rated models, from 18 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - GPQA | 20.47 | Score | 99.3 |
| Open LLM Leaderboard - MuSR | 23.99 | Score | 98.7 |
| Open LLM Leaderboard - BBH | 52.5 | Score | 95.3 |
| Open LLM Leaderboard - MMLU-Pro | 47.54 | Score | 92 |
| GSMA Open-Telco - srsRAN-Bench | 83.36 | Score (%) | 84.9 |
| Open LLM Leaderboard - MATH Level 5 | 29.38 | Score | 81.4 |
| When Context Flips, Safety Breaks | 35.9 | PacifAIst BSR (self-reported) | 63.6 |
| GSMA Open-Telco - TeleTables | 29.93 | Score (%) | 52.3 |
| GSMA Open-Telco LLM Leaderboard | 50.45 | Average Score (%) | 47.7 |
| GSMA Open-Telco - TeleLogs | 19.79 | Score (%) | 45.3 |
| GSMA Open-Telco - TeleMath | 40.07 | Score (%) | 45.3 |
| GSMA Open-Telco - ORAN-Bench | 73.69 | Score (%) | 44.2 |
Interactive version: theaggregate.ai/model?slug=phi-4-14b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.