Mistral 7B: benchmark results
Provider: Mistral. Released 2023-09-27. Access: Open.
Unified ELO 1393 ± 1, rank #1165 of 1392 rated models, from 42 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| The Wrong Kind of Right | 26.8 | MAR Disability (self-reported) | 83.3 |
| T-Eval | 56 | Overall Score (%) | 60 |
| MHGraphBench | 42.1 | Avg$_{All^\ast}$ (self-reported) | 50 |
| RedCode | 61.7 | Score (%) | 50 |
| DiBiMT | 48.31 | Avg Accuracy (%) | 47.4 |
| AutoRace | 26 | Average Score (%, six reasoning tasks) | 44.4 |
| Diagnosing LLM Arbitration Behavior over Pre-e | -19 | Margin (self-reported) | 33.3 |
| LegalCiteBench | 58.5 | Overall (Norm.) (self-reported) | 32.5 |
| ENAMEL | 15.2 | eff@1 (self-reported) | 29 |
| PhysicsFinals | 11.7 | Score (self-reported) | 26.6 |
| CRUXEval | 34.3 | Output Prediction pass@1 (%) | 26.2 |
| AgentBoard | 24.6 | Progress Rate (self-reported) | 25 |
Interactive version: theaggregate.ai/model?slug=mistral-7b · How It Works · Data refreshed daily, snapshot 2026-09-05.