Mistral Medium 3: benchmark results
Mistral's enterprise multimodal mid-tier positioned near Claude 3.7 Sonnet performance at much lower cost, deployable on-prem (May 2025). Provider: Mistral. Released 2025-05-07. Access: API.
Unified ELO 1532 ± 1, rank #531 of 1392 rated models, from 148 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BlueBench - Legal | 67.17 | Score (%) | 100 |
| UGI Leaderboard | 56.45 | UGI Score | 97.1 |
| BlueBench - QA Finance | 33 | Score (%) | 91.2 |
| UGI - Natural Intelligence | 35.26 | NatInt Score | 83.1 |
| BlueBench - Chatbot Abilities | 94.17 | Score (%) | 82.4 |
| BlueBench - Entity Extraction | 74.21 | Score (%) | 82.4 |
| BlueBench - RAG General | 52.57 | Score (%) | 82.4 |
| BlueBench - Reasoning | 75.5 | Score (%) | 79.4 |
| BlueBench | 60.5 | Average Score (%) | 76.5 |
| UGI - Willingness (W/10) | 7.2 | W/10 Score | 73.9 |
| MATH Level 5 | 81.63 | Accuracy (%) | 64.8 |
| Chatbot Arena (Text - Multi-Turn) | 1403 | Arena Score | 64 |
Interactive version: theaggregate.ai/model?slug=mistral-medium-3 · How It Works · Data refreshed daily, snapshot 2026-09-05.