MiniMax M3 (Reasoning): benchmark results
Provider: MiniMax. Released 2026-06-01. Access: Open.
Unified ELO 1609 ± 1, rank #574 of 3078 rated models, from 18 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BALSAM - Program Execution | 94.45 | Overall score (0-100, LLM-judged generation and multiple cho | 92.9 |
| BALSAM - Logic | 38.55 | Overall score (0-100, LLM-judged generation and multiple cho | 82.1 |
| Wolfram LLM Benchmarking Project | 56.1 | Correct Functionality (%) | 81.9 |
| BALSAM - Sequence Tagging | 40.04 | Overall score (0-100, LLM-judged generation and multiple cho | 70.4 |
| BALSAM - Creative Writing | 50.55 | Overall score (0-100, LLM-judged generation and multiple cho | 67.9 |
| BALSAM - Entailment | 76.92 | Overall score (0-100, LLM-judged generation and multiple cho | 57.1 |
| BALSAM - Question Answering | 67.24 | Overall score (0-100, LLM-judged generation and multiple cho | 53.6 |
| BALSAM - Translation/Transliteration | 60.78 | Overall score (0-100, LLM-judged generation and multiple cho | 53.6 |
| BALSAM - Overall | 52.79 | Mean of category overall scores (0-100) | 44.4 |
| BALSAM - Summarization | 49 | Overall score (0-100, LLM-judged generation and multiple cho | 42.9 |
| BALSAM - Reading Comprehension | 49.77 | Overall score (0-100, LLM-judged generation and multiple cho | 35.7 |
| SpeechMap Compliance | 50.5 | % Requests Completed | 33.7 |
Interactive version: theaggregate.ai/model?slug=minimax-m3-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-19.