CatunaMayo-DPO: benchmark results

Provider: Other. Access: Open.

Unified ELO 1504 ± 20, rank #1168 of 2928 rated models, from 12 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard v1 - ARC Challenge72.87Normalized accuracy (%) (25-shot)94
Open LLM Leaderboard v1 - TruthfulQA MC271.82MC2 (%) (0-shot)92.3
Open LLM Leaderboard v1 - GSM8K70.2Accuracy (%) (5-shot)92.1
Open LLM Leaderboard v1 - HellaSwag88.3Normalized accuracy (%) (10-shot)90.8
Open LLM Leaderboard v1 - WinoGrande82.72Accuracy (%) (5-shot)85.2
Open LLM Leaderboard v1 - MMLU65.24Accuracy (%) (5-shot)83.3
Open LLM Leaderboard - MuSR44.5Score80.9
Open LLM Leaderboard - BBH52.24Score60.5
Open LLM Leaderboard - GPQA29.19Score46.9
Open LLM Leaderboard - IFEval42.15Score43.4
Open LLM Leaderboard - MMLU-Pro31.7Score42.8
Open LLM Leaderboard - MATH Level 58.16Score41.4

Interactive version: theaggregate.ai/model?slug=catunamayo-dpo · How It Works · Data refreshed daily, snapshot 2026-09-23.