DiscoLM-70B: benchmark results
Provider: Other. Released 2023-11-19. Access: Open.
Unified ELO 1488 ± 20, rank #1369 of 2928 rated models, from 108 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard v1 - MMLU | 68.58 | Accuracy (%) (5-shot) | 91.6 |
| Open LLM Leaderboard v1 - WinoGrande | 83.58 | Accuracy (%) (5-shot) | 89.7 |
| Open LLM Leaderboard v1 - ARC Challenge | 68.77 | Normalized accuracy (%) (25-shot) | 81.4 |
| Open LLM Leaderboard v1 - HellaSwag | 86.1 | Normalized accuracy (%) (10-shot) | 78.7 |
| Open LLM Leaderboard v1 - GSM8K | 63.53 | Accuracy (%) (5-shot) | 77 |
| EuroEval Portuguese NLU - HAREM | 49.95 | Named entity recognition Score (%) | 76.8 |
| EuroEval Italian NLU - MultiNERD IT | 72.48 | Named entity recognition Score (%) | 75.1 |
| EuroEval Faroese NLU - FONE | 71.63 | Named entity recognition Score (%) | 72.8 |
| Open LLM Leaderboard v1 - TruthfulQA MC2 | 57.64 | MC2 (%) (0-shot) | 70.6 |
| EuroEval German NLU - GermEval | 63.26 | Named entity recognition Score (%) | 69.6 |
| EuroEval Danish NLU - Dansk | 57.93 | Named entity recognition Score (%) | 68.4 |
| EuroEval Spanish NLU - CoNLL ES | 65.78 | Named entity recognition Score (%) | 64.9 |
Interactive version: theaggregate.ai/model?slug=discolm-70b · How It Works · Data refreshed daily, snapshot 2026-09-23.