Magistral Small: benchmark results
Mistral's Apache-2.0 open 24B reasoning model, SFT-distilled from Magistral Medium traces plus RL (June 2025). Provider: Mistral. Released 2025-06-10. Access: Open.
Unified ELO 1477 ± 1, rank #798 of 1392 rated models, from 112 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Medmarks - LongHealth Task 2 | 88.96 | Score (%) | 90 |
| Greek MMLU - Greek-Specific | 78.83 | Accuracy (%) | 81.5 |
| Greek MMLU | 71.27 | Accuracy (%) | 80.2 |
| BRIDGE Medical Leaderboard - Few-Shot | 49.15 | Average Performance (%) | 79.6 |
| Greek MMLU - Humanities | 73.15 | Accuracy (%) | 79 |
| Greek MMLU - STEM | 69.06 | Accuracy (%) | 77.8 |
| Greek MMLU - Social Sciences | 74.61 | Accuracy (%) | 77.8 |
| SpeechMap Compliance | 78.1 | % Requests Completed | 74.9 |
| Greek MMLU - Other | 65.94 | Accuracy (%) | 72.8 |
| BRIDGE Medical Leaderboard | 40.29 | Average Performance (%) | 65.7 |
| Medmarks - PubHealthBench Reviewed | 85.04 | Score (%) | 65 |
| BRIDGE Medical Leaderboard - Zero-Shot | 37.56 | Average Performance (%) | 61.1 |
Interactive version: theaggregate.ai/model?slug=magistral-small · How It Works · Data refreshed daily, snapshot 2026-09-05.