Mistral Large 3: benchmark results
Mistral AI's Large 3, a 675B sparse MoE (41B active) flagship (December 2025). Provider: Mistral. Released 2025-12-02. Access: API.
Unified ELO 1563 ± 1, rank #391 of 1392 rated models, from 320 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Stats (MM-MT-Bench) | 84.9 | Score (%) | 100 |
| ROK-FORTRESS | 28.4 | Δ_ling (pp) (self-reported) | 100 |
| SpeechMap Compliance | 98.2 | % Requests Completed | 99.1 |
| UGI Leaderboard | 56.86 | UGI Score | 97.7 |
| AGC-Bench - ocw | 1.63 | Dataset z-score | 97.6 |
| StereoTales | 109 | Emissions (self-reported) | 95.5 |
| AGC-Bench - cpers | 1.14 | Dataset z-score | 93.9 |
| AGC-Bench - mops | 1.46 | Dataset z-score | 93.9 |
| LLM Stats (Wild Bench) | 68.5 | Score (%) | 92.9 |
| AA Omniscience - Software Engineering (SWE) - Java | 30 | Accuracy (%) | 90.4 |
| AGC-Bench - thenextchapter | 1.13 | Dataset z-score | 90.2 |
| AA Omniscience - Software Engineering (SWE) - Kotlin | 36 | Accuracy (%) | 89.4 |
Interactive version: theaggregate.ai/model?slug=mistral-large-3 · How It Works · Data refreshed daily, snapshot 2026-09-05.