Phi-3.5-mini-instruct: benchmark results

Microsoft's 3.8B MIT-licensed small model with a 128K context and improved multilingual coverage. Provider: Microsoft. Released 2024-08-20. Access: Open.

Unified ELO 1424 ± 1, rank #1052 of 1392 rated models, from 136 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Stats (SQuALITY)24.3Score (%)100
LIBRA - MatreshkaYesNo73.42Dataset Total Score (%)91.7
Open Japanese LLM - Xlsum JA Bert Score JA F170.67Score (%)88.8
LLM Stats (Social IQa)74.7Score (%)87.5
GSM8K86.2Accuracy (%)86
Open LLM Leaderboard - GPQA11.97Score84.3
Open Japanese LLM - Xlsum JA Rouge129.95Score (%)80.8
Open Japanese LLM - Xlsum JA RougeLsum25.34Score (%)80.3
Enkrypt AI - Jailbreak Risk4.28Risk Score79.8
Open Japanese LLM - SUM10.73Score (%)76.6
Open Japanese LLM - Xlsum JA Rouge210.73Score (%)76.6
Open LLM Leaderboard - BBH36.75Score76

Interactive version: theaggregate.ai/model?slug=phi-3-5-mini-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.