Phi-3-small-8k-instruct — benchmark results

Microsoft's 7B Phi-3-small instruct model (May 2024) with 8K context and blocksparse attention, trained on 4.8T tokens to beat GPT-3.5 Turbo. Provider: Microsoft. Released 2024-05-21. Access: Open.

Unified ELO 1450 ± 16, rank #999 of 1776 rated models, from 48 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
OpenBookQA88Accuracy (%)98.8
ARC Challenge (AI2)90.7Accuracy (%)92.3
Big-Bench Hard79.1Average (%)89.6
ChineseSafe Benchmark72.73Accuracy (%)88.9
Open LLM Leaderboard - MuSR16.77Score86.8
Epoch AI - Adversarial Nli58.1Score86.4
Open LLM Leaderboard - BBH46.21Score86.1
CanAiCode100Junior-v2 Python Pass Rate (%)85.6
LiveBench Zebra Puzzle40Score85.4
Open CoT - LSAT Reading Comprehension19.7CoT Gain (%)84
Open LLM Leaderboard - MMLU-Pro38.96Score83.8
Open CoT - LSAT Logical Reasoning17.84CoT Gain (%)83.6

Interactive version: theaggregate.ai/model?slug=phi-3-small-8k-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.