Phi-3 Mini 4K Instruct — benchmark results

Provider: Microsoft. Released 2024-04-23. Access: Open.

Unified ELO 1393 ± 15, rank #1259 of 1776 rated models, from 90 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HumanLikeness - Discourse-179.7Humanlike Score (%)100
HumanLikeness - Syntax-189.36Humanlike Score (%)100
OpenBookQA88Accuracy (%)98.8
HumanLikeness - Discourse-276.18Humanlike Score (%)94.7
HumanLikeness - Sound-267.57Humanlike Score (%)94.7
HumanLikeness - Meaning-176.07Humanlike Score (%)89.5
HumanLikeness - Overall64.73Overall Humanlike (%)89.5
BiGGen-Bench3.82Average Score (1-5)84.3
Big-Bench Hard71.7Average (%)82.3
ARC Challenge (AI2)84.9Accuracy (%)82.1
Open LLM Leaderboard - GPQA10.96Score81.4
Open CoT - LSAT Analytical Reasoning6.52CoT Gain (%)81.3

Interactive version: theaggregate.ai/model?slug=phi-3-mini-4k-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.