MiniMax-M3 — benchmark results
MiniMax M3 model for reasoning and agentic tasks. Provider: MiniMax. Released 2026-06-01. Access: Open.
Unified ELO 1779 ± 13, rank #163 of 1776 rated models, from 165 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Stats (OmniDocBench 1.5) | 91.6 | Score (%) | 100 |
| LLM Stats (PostTrainBench) | 37.1 | Score (%) | 100 |
| WorldCupBench | 87 | Quiniela Points | 100 |
| AA IFBench | 82.86 | Accuracy (%) | 99.6 |
| AA GPQA Diamond | 92.93 | Accuracy (%) | 98.7 |
| AA Long Context Reasoning | 74 | Accuracy (%) | 98.3 |
| MedScribe | 87.25 | Score (self-reported) | 96.2 |
| AA Humanity's Last Exam | 37.12 | Accuracy (%) | 94.8 |
| Vals AI GPQA | 92.68 | Accuracy (%) | 94.7 |
| Artificial Analysis Intelligence Index | 44.44 | Intelligence Index | 94.5 |
| Vals AI MedScribe | 87.25 | Accuracy (%) | 94.4 |
| Vals AI CorpFin v2 | 68.1 | Accuracy (%) | 94.3 |
Interactive version: theaggregate.ai/model?slug=minimax-m3 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.