Darkest-muse-v1 — benchmark results
sam-paech's creative-writing merge of two Gemma 2 9B fine-tunes (Quill and Delirium) DPO-trained on Gutenberg fiction extracts. Provider: Other. Released 2024-10-22. Access: Open.
Unified ELO 1501 ± 29, rank #796 of 1776 rated models, from 12 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - IFEval | 73.44 | Score | 87.9 |
| Open LLM Leaderboard - GPQA | 12.53 | Score | 85.8 |
| Open LLM Leaderboard - MuSR | 15.28 | Score | 82.3 |
| Open LLM Leaderboard - BBH | 42.61 | Score | 81.7 |
| Open LLM Leaderboard - MMLU-Pro | 35.38 | Score | 76.1 |
| Open LLM Leaderboard - MATH Level 5 | 21.45 | Score | 73.5 |
| Creative Writing v3 | 1173 | Elo score (self-reported) | 28.7 |
| EQ-Bench Creative Writing v3 | 1027.2 | Elo | 27.9 |
| UGI - Writing | 20.03 | Writing Score | 13.6 |
| UGI - Natural Intelligence | 13.17 | NatInt Score | 12.1 |
| UGI - Willingness (W/10) | 1.5 | W/10 Score | 5.6 |
| UGI Leaderboard | 11.42 | UGI Score | 3.2 |
Interactive version: theaggregate.ai/model?slug=darkest-muse-v1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.