Phi-3-small-8k-instruct — benchmark results
Microsoft's 7B Phi-3-small instruct model (May 2024) with 8K context and blocksparse attention, trained on 4.8T tokens to beat GPT-3.5 Turbo. Provider: Microsoft. Released 2024-05-21. Access: Open.
Unified ELO 1450 ± 16, rank #999 of 1776 rated models, from 48 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| OpenBookQA | 88 | Accuracy (%) | 98.8 |
| ARC Challenge (AI2) | 90.7 | Accuracy (%) | 92.3 |
| Big-Bench Hard | 79.1 | Average (%) | 89.6 |
| ChineseSafe Benchmark | 72.73 | Accuracy (%) | 88.9 |
| Open LLM Leaderboard - MuSR | 16.77 | Score | 86.8 |
| Epoch AI - Adversarial Nli | 58.1 | Score | 86.4 |
| Open LLM Leaderboard - BBH | 46.21 | Score | 86.1 |
| CanAiCode | 100 | Junior-v2 Python Pass Rate (%) | 85.6 |
| LiveBench Zebra Puzzle | 40 | Score | 85.4 |
| Open CoT - LSAT Reading Comprehension | 19.7 | CoT Gain (%) | 84 |
| Open LLM Leaderboard - MMLU-Pro | 38.96 | Score | 83.8 |
| Open CoT - LSAT Logical Reasoning | 17.84 | CoT Gain (%) | 83.6 |
Interactive version: theaggregate.ai/model?slug=phi-3-small-8k-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.