Mahou-1.2a-llama3-8B: benchmark results

Provider: Other. Access: Open.

Unified ELO 1498 ± 19, rank #1252 of 2928 rated models, from 12 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard v1 - GSM8K71.11Accuracy (%) (5-shot)95.2
Open LLM Leaderboard v1 - MMLU68.52Accuracy (%) (5-shot)91.5
Open LLM Leaderboard v1 - ARC Challenge68.86Normalized accuracy (%) (25-shot)81.6
Open LLM Leaderboard v1 - TruthfulQA MC258.76MC2 (%) (0-shot)72.6
Open LLM Leaderboard - MMLU-Pro38.17Score66.3
Open LLM Leaderboard v1 - HellaSwag83.97Normalized accuracy (%) (10-shot)62.8
Open LLM Leaderboard - IFEval50.93Score60.8
Open LLM Leaderboard v1 - WinoGrande77.58Accuracy (%) (5-shot)57
Open LLM Leaderboard - BBH50.94Score53.3
Open LLM Leaderboard - GPQA28.86Score43.2
Open LLM Leaderboard - MATH Level 58.38Score42.2
Open LLM Leaderboard - MuSR38.47Score32.3

Interactive version: theaggregate.ai/model?slug=mahou-1-2a-llama3-8b · How It Works · Data refreshed daily, snapshot 2026-09-23.