Dolly V2 12B — benchmark results
Provider: Databricks. Released 2023-04-12. Access: Open.
Unified ELO 1102 ± 35, rank #1760 of 1776 rated models, from 92 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLMsPark - Dictator Game | 1036.5 | Elo Rating | 92.3 |
| LLMsPark - Nim Game | 1180.3 | Elo Rating | 92.3 |
| MMLU-by-task - Global Facts | 39 | Accuracy (%) | 85.5 |
| MMLU-by-task - Machine Learning | 34.82 | Accuracy (%) | 60.6 |
| LLMsPark - Trust Game | 1018.7 | Elo Rating | 53.8 |
| MMLU-by-task - High School Mathematics | 26.67 | Accuracy (%) | 45.8 |
| ToolBench - WebShop Long | 0 | Task Score | 44.2 |
| MMLU-by-task - College Physics | 23.53 | Accuracy (%) | 41.5 |
| MMLU-by-task - Elementary Mathematics | 26.98 | Accuracy (%) | 39 |
| MMLU-by-task - HellaSwag | 54.63 | Accuracy (%) | 38.6 |
| MMLU-by-task - High School Computer Science | 36 | Accuracy (%) | 38.2 |
| ToolBench - Tabletop | 7.6 | Task Score | 33.7 |
Interactive version: theaggregate.ai/model?slug=dolly-v2-12b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.