Mistral Medium 3 — benchmark results
Mistral's enterprise multimodal mid-tier positioned near Claude 3.7 Sonnet performance at much lower cost, deployable on-prem (May 2025). Provider: Mistral. Released 2025-05-07. Access: API.
Unified ELO 1553 ± 12, rank #612 of 1776 rated models, from 121 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BlueBench - Legal | 67.17 | Score (%) | 100 |
| UGI Leaderboard | 56.45 | UGI Score | 97.1 |
| BlueBench - QA Finance | 33 | Score (%) | 91.2 |
| UGI - Natural Intelligence | 35.26 | NatInt Score | 84 |
| BlueBench - Chatbot Abilities | 94.17 | Score (%) | 82.4 |
| BlueBench - Entity Extraction | 74.21 | Score (%) | 82.4 |
| BlueBench - RAG General | 52.57 | Score (%) | 82.4 |
| BlueBench - Reasoning | 75.5 | Score (%) | 79.4 |
| BlueBench | 60.5 | Average Score (%) | 76.5 |
| YapBench | 361 | YapIndex (lower is better) | 73.8 |
| UGI - Willingness (W/10) | 7.2 | W/10 Score | 73.6 |
| AA MATH-500 | 90.67 | Accuracy (%) | 66.8 |
Interactive version: theaggregate.ai/model?slug=mistral-medium-3 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.