Mistral Large: benchmark results
Mistral AI's first-generation Large, an API-only flagship (February 2024). Provider: Mistral. Released 2024-02-26. Access: API.
Unified ELO 1509 ± 1, rank #653 of 1392 rated models, from 94 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BlueBench - Safety | 88.04 | Score (%) | 94.1 |
| InfiBench | 58.22 | Score (%) | 92.4 |
| BiGGen-Bench | 3.93 | Average Score (1-5) | 88.2 |
| MAGI-Hard | 67.69 | Accuracy (%, 2024-05 snapshot) | 86.2 |
| SEAL - Adversarial Robustness | 37 | Score | 85.7 |
| BlueBench - Bias | 96.97 | Score (%) | 82.4 |
| BlueBench - Summarization | 18.24 | Score (%) | 82.4 |
| CZ-EVAL - Analytical | 38.52 | Accuracy (%) | 81.8 |
| CZ-EVAL - Culture | 85.94 | Accuracy (%) | 81.8 |
| ProLLM - Q&A Assistant | 96.4 | Score (%) | 80.4 |
| BlueBench - Knowledge | 55.1 | Score (%) | 79.4 |
| CZ-EVAL - Critical Thinking | 66.78 | Accuracy (%) | 72.7 |
Interactive version: theaggregate.ai/model?slug=mistral-large · How It Works · Data refreshed daily, snapshot 2026-09-05.