Mistral Medium 3: benchmark results

Mistral's enterprise multimodal mid-tier positioned near Claude 3.7 Sonnet performance at much lower cost, deployable on-prem (May 2025). Provider: Mistral. Released 2025-05-07. Access: API.

Unified ELO 1532 ± 1, rank #531 of 1392 rated models, from 148 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BlueBench - Legal67.17Score (%)100
UGI Leaderboard56.45UGI Score97.1
BlueBench - QA Finance33Score (%)91.2
UGI - Natural Intelligence35.26NatInt Score83.1
BlueBench - Chatbot Abilities94.17Score (%)82.4
BlueBench - Entity Extraction74.21Score (%)82.4
BlueBench - RAG General52.57Score (%)82.4
BlueBench - Reasoning75.5Score (%)79.4
BlueBench60.5Average Score (%)76.5
UGI - Willingness (W/10)7.2W/10 Score73.9
MATH Level 581.63Accuracy (%)64.8
Chatbot Arena (Text - Multi-Turn)1403Arena Score64

Interactive version: theaggregate.ai/model?slug=mistral-medium-3 · How It Works · Data refreshed daily, snapshot 2026-09-05.