magnum-v2-12B — benchmark results

Anthracite's magnum v2 fine-tune of Mistral Nemo Base 2407 (12B), aimed at replicating the prose quality of Claude 3 Sonnet and Opus. Provider: Other. Released 2024-08-03. Access: Open.

Unified ELO 1440 ± 29, rank #1043 of 1776 rated models, from 10 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard - MuSR11.37Score59.4
UGI - Writing32.97Writing Score48.3
Open LLM Leaderboard - BBH28.79Score47.6
Open LLM Leaderboard - GPQA5.48Score46
Open LLM Leaderboard - MMLU-Pro24.08Score42.8
Open LLM Leaderboard - IFEval37.62Score36.7
Open LLM Leaderboard - MATH Level 55.44Score30.2
UGI - Willingness (W/10)3.2W/10 Score24.1
UGI - Natural Intelligence15.74NatInt Score22.5
UGI Leaderboard21.54UGI Score14.2

Interactive version: theaggregate.ai/model?slug=magnum-v2-12b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.