MAI-Thinking-1 — benchmark results
Microsoft AI's first in-house reasoning model, a sparse MoE (1T total, 35B active). Provider: Microsoft. Released 2026-06-02. Access: API.
Unified ELO 1642 ± 38, rank #403 of 1841 rated models, from 8 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Stats (AIME 2026) | 94.5 | Score (%) | 75 |
| ZeroEval GPQA Diamond | 84.2 | GPQA Diamond Score | 74.1 |
| LLM Stats (LongBench v2) | 61 | Score (%) | 66.7 |
| LLM Stats (Multi-Challenge) | 53 | Score (%) | 46.4 |
| MedXpertQA | 43 | Score (self-reported) | 44 |
| BenchLM | 50.4 | Overall Score | 42.6 |
| LLM Stats (HMMT Feb 26) | 84.9 | Score (%) | 21.4 |
| LLM Stats (HealthBench Professional) | 35 | Score (%) | 0 |
Interactive version: theaggregate.ai/model?slug=mai-thinking-1 · How It Works · Data refreshed daily, snapshot 2026-07-25.