Magistral Small: benchmark results

Mistral's Apache-2.0 open 24B reasoning model, SFT-distilled from Magistral Medium traces plus RL (June 2025). Provider: Mistral. Released 2025-06-10. Access: Open.

Unified ELO 1477 ± 1, rank #798 of 1392 rated models, from 112 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Medmarks - LongHealth Task 288.96Score (%)90
Greek MMLU - Greek-Specific78.83Accuracy (%)81.5
Greek MMLU71.27Accuracy (%)80.2
BRIDGE Medical Leaderboard - Few-Shot49.15Average Performance (%)79.6
Greek MMLU - Humanities73.15Accuracy (%)79
Greek MMLU - STEM69.06Accuracy (%)77.8
Greek MMLU - Social Sciences74.61Accuracy (%)77.8
SpeechMap Compliance78.1% Requests Completed74.9
Greek MMLU - Other65.94Accuracy (%)72.8
BRIDGE Medical Leaderboard40.29Average Performance (%)65.7
Medmarks - PubHealthBench Reviewed85.04Score (%)65
BRIDGE Medical Leaderboard - Zero-Shot37.56Average Performance (%)61.1

Interactive version: theaggregate.ai/model?slug=magistral-small · How It Works · Data refreshed daily, snapshot 2026-09-05.