Phi-lthy4 — benchmark results

SicariusSicariiStuff's pruned 11.9B rebuild of Phi-4, re-architected and merged for roleplay and creative writing with low refusals. Provider: Microsoft. Released 2025-02-12. Access: Open.

Unified ELO 1485 ± 36, rank #859 of 1776 rated models, from 9 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard - IFEval76.79Score92.6
Open LLM Leaderboard - BBH40.15Score79.3
Open LLM Leaderboard - MMLU-Pro37.04Score78.6
Open LLM Leaderboard - MATH Level 513.67Score59.5
UGI - Willingness (W/10)5.8W/10 Score49.3
Open LLM Leaderboard - MuSR9.04Score43.7
Open LLM Leaderboard - GPQA4.92Score41.7
UGI Leaderboard27.24UGI Score24.2
UGI - Natural Intelligence14.56NatInt Score16.7

Interactive version: theaggregate.ai/model?slug=phi-lthy4 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.