Phi-lthy4: benchmark results

SicariusSicariiStuff's pruned 11.9B rebuild of Phi-4, re-architected and merged for roleplay and creative writing with low refusals. Provider: Microsoft. Released 2025-02-12. Access: Open.

Unified ELO 1514 ± 1, rank #630 of 1392 rated models, from 9 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard - IFEval76.79Score92.6
Open LLM Leaderboard - BBH40.15Score79.2
Open LLM Leaderboard - MMLU-Pro37.04Score78.6
Open LLM Leaderboard - MATH Level 513.67Score59.5
UGI - Willingness (W/10)5.8W/10 Score49.8
Open LLM Leaderboard - MuSR9.04Score43.7
Open LLM Leaderboard - GPQA4.92Score41.8
UGI Leaderboard27.24UGI Score24
UGI - Natural Intelligence14.56NatInt Score16.5

Interactive version: theaggregate.ai/model?slug=phi-lthy4 · How It Works · Data refreshed daily, snapshot 2026-09-05.