Mistral Large: benchmark results

Mistral AI's first-generation Large, an API-only flagship (February 2024). Provider: Mistral. Released 2024-02-26. Access: API.

Unified ELO 1509 ± 1, rank #653 of 1392 rated models, from 94 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BlueBench - Safety88.04Score (%)94.1
InfiBench58.22Score (%)92.4
BiGGen-Bench3.93Average Score (1-5)88.2
MAGI-Hard67.69Accuracy (%, 2024-05 snapshot)86.2
SEAL - Adversarial Robustness37Score85.7
BlueBench - Bias96.97Score (%)82.4
BlueBench - Summarization18.24Score (%)82.4
CZ-EVAL - Analytical38.52Accuracy (%)81.8
CZ-EVAL - Culture85.94Accuracy (%)81.8
ProLLM - Q&A Assistant96.4Score (%)80.4
BlueBench - Knowledge55.1Score (%)79.4
CZ-EVAL - Critical Thinking66.78Accuracy (%)72.7

Interactive version: theaggregate.ai/model?slug=mistral-large · How It Works · Data refreshed daily, snapshot 2026-09-05.