Mistral-7B-v0.1-flashback-v2: benchmark results
Provider: Mistral. Released 2023-12-04. Access: Open.
Unified ELO 1428 ± 20, rank #1913 of 2928 rated models, from 32 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard v1 - WinoGrande | 77.19 | Accuracy (%) (5-shot) | 54 |
| Open LLM Leaderboard v1 - GSM8K | 29.42 | Accuracy (%) (5-shot) | 46.7 |
| Open LLM Leaderboard v1 - MMLU | 59.98 | Accuracy (%) (5-shot) | 46.4 |
| Open LLM Leaderboard v1 - HellaSwag | 80.74 | Normalized accuracy (%) (10-shot) | 42.2 |
| EuroEval Italian NLU - MultiNERD IT | 56.62 | Named entity recognition Score (%) | 41.9 |
| Open LLM Leaderboard v1 - ARC Challenge | 57.17 | Normalized accuracy (%) (25-shot) | 36.8 |
| EuroEval Dutch NLU - DBRD | 74.48 | Sentiment classification Score (%) | 30.5 |
| EuroEval Danish Knowledge | 21.47 | Knowledge Average Score (%) | 29.4 |
| EuroEval Spanish NLU - ScaLA ES | 1.23 | Linguistic acceptability Score (%) | 28.3 |
| EuroEval Danish Knowledge - Danish Citizen Tests | 21.47 | MCC (x100) | 26.7 |
| EuroEval Spanish NLU - CoNLL ES | 48.21 | Named entity recognition Score (%) | 26.6 |
| EuroEval English Knowledge | 24 | Knowledge Average Score (%) | 20.5 |
Interactive version: theaggregate.ai/model?slug=mistral-7b-v0-1-flashback-v2 · How It Works · Data refreshed daily, snapshot 2026-09-23.