Mistral Large 3 — benchmark results

Mistral Large 3 model row. Provider: Mistral. Released 2025-12-02. Access: API.

Unified ELO 1606 ± 12, rank #466 of 1776 rated models, from 274 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Stats (MM-MT-Bench)84.9Score (%)100
ROK-FORTRESS28.4Δ_ling (pp) (self-reported)100
XL-SafetyBench98.8Overall ASR (self-reported)100
SpeechMap Compliance98.2% Requests Completed98.9
UGI Leaderboard56.86UGI Score97.7
AGC-Bench - ocw1.63Dataset z-score97.6
AGC-Bench - cpers1.14Dataset z-score93.9
AGC-Bench - mops1.46Dataset z-score93.9
LLM Stats (Wild Bench)68.5Score (%)92.9
AGC-Bench - thenextchapter1.13Dataset z-score90.2
UGI - Natural Intelligence39.98NatInt Score87.2
HumainE Leaderboard32.68HumainE Score86.8

Interactive version: theaggregate.ai/model?slug=mistral-large-3 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.