dolphin-2.9.4-llama3.1-8B — benchmark results
Provider: Cognitive Computations. Released 2024-08-04. Access: Open.
Unified ELO 1137 ± 58, rank #1749 of 1776 rated models, from 7 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - IFEval | 27.57 | Score | 24.6 |
| Open LLM Leaderboard - BBH | 8.97 | Score | 18.8 |
| Open LLM Leaderboard - GPQA | 1.79 | Score | 17.3 |
| Open LLM Leaderboard - MMLU-Pro | 2.63 | Score | 10.2 |
| Open LLM Leaderboard - MATH Level 5 | 1.21 | Score | 8.7 |
| Open Korean LLM Leaderboard | 28.07 | Average Score (%) | 5.6 |
| Open LLM Leaderboard - MuSR | 0.62 | Score | 0.4 |
Interactive version: theaggregate.ai/model?slug=dolphin-2-9-4-llama3-1-8b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.