MiniMax M3 (Reasoning): benchmark results

Provider: MiniMax. Released 2026-06-01. Access: Open.

Unified ELO 1609 ± 1, rank #574 of 3078 rated models, from 18 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BALSAM - Program Execution94.45Overall score (0-100, LLM-judged generation and multiple cho92.9
BALSAM - Logic38.55Overall score (0-100, LLM-judged generation and multiple cho82.1
Wolfram LLM Benchmarking Project56.1Correct Functionality (%)81.9
BALSAM - Sequence Tagging40.04Overall score (0-100, LLM-judged generation and multiple cho70.4
BALSAM - Creative Writing50.55Overall score (0-100, LLM-judged generation and multiple cho67.9
BALSAM - Entailment76.92Overall score (0-100, LLM-judged generation and multiple cho57.1
BALSAM - Question Answering67.24Overall score (0-100, LLM-judged generation and multiple cho53.6
BALSAM - Translation/Transliteration60.78Overall score (0-100, LLM-judged generation and multiple cho53.6
BALSAM - Overall52.79Mean of category overall scores (0-100)44.4
BALSAM - Summarization49Overall score (0-100, LLM-judged generation and multiple cho42.9
BALSAM - Reading Comprehension49.77Overall score (0-100, LLM-judged generation and multiple cho35.7
SpeechMap Compliance50.5% Requests Completed33.7

Interactive version: theaggregate.ai/model?slug=minimax-m3-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-19.