Phi-1.5 — benchmark results
Provider: Microsoft. Released 2023-09-11. Access: Open.
Unified ELO 1168 ± 23, rank #1720 of 1776 rated models, from 123 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MMLU-by-task - College Mathematics | 41 | Accuracy (%) | 97.3 |
| MMLU-by-task - Abstract Algebra | 34 | Accuracy (%) | 82.1 |
| AI Energy Score (Text Generation) | 5 | Energy Score (1-5) | 79.5 |
| MMLU-by-task - College Computer Science | 46 | Accuracy (%) | 75.4 |
| MMLU-by-task - Machine Learning | 38.39 | Accuracy (%) | 75.4 |
| MMLU-by-task - Electrical Engineering | 48.28 | Accuracy (%) | 68.4 |
| MMLU-by-task - College Physics | 27.45 | Accuracy (%) | 62.9 |
| MMLU-by-task - Business Ethics | 52 | Accuracy (%) | 61.7 |
| MMLU-by-task - Moral Disputes | 54.91 | Accuracy (%) | 61.5 |
| MMLU-by-task - Elementary Mathematics | 30.16 | Accuracy (%) | 57.8 |
| MMLU-by-task - Sociology | 64.68 | Accuracy (%) | 56.8 |
| MMLU-by-task - High School Microeconomics | 45.38 | Accuracy (%) | 56.1 |
Interactive version: theaggregate.ai/model?slug=phi-1-5 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.