Phi-lthy4 — benchmark results
SicariusSicariiStuff's pruned 11.9B rebuild of Phi-4, re-architected and merged for roleplay and creative writing with low refusals. Provider: Microsoft. Released 2025-02-12. Access: Open.
Unified ELO 1485 ± 36, rank #859 of 1776 rated models, from 9 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - IFEval | 76.79 | Score | 92.6 |
| Open LLM Leaderboard - BBH | 40.15 | Score | 79.3 |
| Open LLM Leaderboard - MMLU-Pro | 37.04 | Score | 78.6 |
| Open LLM Leaderboard - MATH Level 5 | 13.67 | Score | 59.5 |
| UGI - Willingness (W/10) | 5.8 | W/10 Score | 49.3 |
| Open LLM Leaderboard - MuSR | 9.04 | Score | 43.7 |
| Open LLM Leaderboard - GPQA | 4.92 | Score | 41.7 |
| UGI Leaderboard | 27.24 | UGI Score | 24.2 |
| UGI - Natural Intelligence | 14.56 | NatInt Score | 16.7 |
Interactive version: theaggregate.ai/model?slug=phi-lthy4 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.