Phi-3-mini-4K-instruct-cpo-simpo: benchmark results

Community CPO-SimPO preference-tuned variant of Microsoft's Phi-3 Mini 4K Instruct (3.8B), reported to improve GSM8K and TruthfulQA. Provider: Microsoft. Released 2024-06-24. Access: Open.

Unified ELO 1436 ± 1, rank #1006 of 1392 rated models, from 130 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard - GPQA10.74Score80.7
Open LLM Leaderboard - BBH39.15Score78.7
Open Arabic LLM - Arabic MMLU HT Abstract Algebra33Accuracy (%)70.4
Open LLM Leaderboard - IFEval57.14Score69
Open LLM Leaderboard - MMLU-Pro31.78Score68.8
Open Arabic LLM - Aratrust Offensive84.06Accuracy (%)63.4
Open LLM Leaderboard - MATH Level 515.71Score63.3
Open Arabic LLM - Arabic MMLU HT College Mathematics33Accuracy (%)60.2
Open Arabic LLM - Arabic MMLU HT High School Mathematics30.37Accuracy (%)53.7
Open Arabic LLM - Alghafa Multiple Choice Rating Sentiment NO Neutral Task76.34Accuracy (%)51.9
Open Arabic LLM - Alghafa Multiple Choice Rating Sentiment Task51.03Accuracy (%)46.3
Open Arabic LLM - Arabic MMLU HT Medical Genetics44Accuracy (%)43.8

Interactive version: theaggregate.ai/model?slug=phi-3-mini-4k-instruct-cpo-simpo · How It Works · Data refreshed daily, snapshot 2026-09-05.