Mistral Medium 3.1 — benchmark results
Mistral's Medium 3.1 frontier multimodal model (August 2025). Provider: Mistral. Released 2025-08-13. Access: API.
Unified ELO 1625 ± 11, rank #423 of 1776 rated models, from 404 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - cpers | 1.65 | Dataset z-score | 100 |
| AGC-Bench - mops | 2.56 | Dataset z-score | 100 |
| AGC-Bench - story_quality | 1.38 | Dataset z-score | 100 |
| AGC-Bench - pun_eval | 1.42 | Dataset z-score | 98.8 |
| AGC-Bench - tinyfabulist | 1.83 | Dataset z-score | 98.8 |
| AGC-Bench - thenextchapter | 2.08 | Dataset z-score | 97.6 |
| AGC-Bench - tinystories | 1.6 | Dataset z-score | 96.3 |
| AGC-Bench - unfun_corpus | 1.03 | Dataset z-score | 96.2 |
| Galileo Agent - Insurance Accuracy | 70 | Accuracy (%) | 95.2 |
| Galileo Agent - Investment Accuracy | 57 | Accuracy (%) | 95.2 |
| Galileo Agent - Telecom Accuracy | 59 | Accuracy (%) | 95.2 |
| Galileo Agent Leaderboard | 61 | Avg Accuracy (%) | 95.2 |
Interactive version: theaggregate.ai/model?slug=mistral-medium-3-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.