magnum-v4-22B — benchmark results

Anthracite's community magnum v4 fine-tune of Mistral Small Instruct 2409 (22B) aiming to replicate Claude 3 Sonnet/Opus prose quality. Provider: Other. Released 2024-10-20. Access: Open.

Unified ELO 1526 ± 15, rank #700 of 1776 rated models, from 10 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard - GPQA10.4Score79.6
Open LLM Leaderboard - MuSR13.43Score74
Open LLM Leaderboard - BBH35.55Score71.9
Open LLM Leaderboard - MATH Level 520.02Score70.9
Open LLM Leaderboard - IFEval56.29Score68.2
Open LLM Leaderboard - MMLU-Pro31.44Score67.1
UGI - Writing35.22Writing Score56.4
UGI - Natural Intelligence21.94NatInt Score51.4
UGI - Willingness (W/10)4.8W/10 Score36.1
UGI Leaderboard27.33UGI Score24.3

Interactive version: theaggregate.ai/model?slug=magnum-v4-22b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.