Dolly V2 12B — benchmark results

Provider: Databricks. Released 2023-04-12. Access: Open.

Unified ELO 1102 ± 35, rank #1760 of 1776 rated models, from 92 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLMsPark - Dictator Game1036.5Elo Rating92.3
LLMsPark - Nim Game1180.3Elo Rating92.3
MMLU-by-task - Global Facts39Accuracy (%)85.5
MMLU-by-task - Machine Learning34.82Accuracy (%)60.6
LLMsPark - Trust Game1018.7Elo Rating53.8
MMLU-by-task - High School Mathematics26.67Accuracy (%)45.8
ToolBench - WebShop Long0Task Score44.2
MMLU-by-task - College Physics23.53Accuracy (%)41.5
MMLU-by-task - Elementary Mathematics26.98Accuracy (%)39
MMLU-by-task - HellaSwag54.63Accuracy (%)38.6
MMLU-by-task - High School Computer Science36Accuracy (%)38.2
ToolBench - Tabletop7.6Task Score33.7

Interactive version: theaggregate.ai/model?slug=dolly-v2-12b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.