pygmalion-6B: benchmark results

Provider: Other. Access: Open.

Unified ELO 1243 ± 20, rank #2819 of 2928 rated models, from 12 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard - MuSR36.84Score22.9
Open LLM Leaderboard v1 - GSM8K2.05Accuracy (%) (5-shot)19.2
Open LLM Leaderboard v1 - HellaSwag67.47Normalized accuracy (%) (10-shot)18.5
Open LLM Leaderboard v1 - ARC Challenge40.53Normalized accuracy (%) (25-shot)16.8
Open LLM Leaderboard v1 - WinoGrande62.51Accuracy (%) (5-shot)15.5
Open LLM Leaderboard - IFEval20.91Score13.2
Open LLM Leaderboard - BBH31.99Score11.9
Open LLM Leaderboard v1 - MMLU25.73Accuracy (%) (5-shot)9.6
Open LLM Leaderboard - MMLU-Pro11.84Score9.1
Open LLM Leaderboard - MATH Level 50.83Score6
Open LLM Leaderboard - GPQA24.92Score4.6
Open LLM Leaderboard v1 - TruthfulQA MC232.53MC2 (%) (0-shot)0.2

Interactive version: theaggregate.ai/model?slug=pygmalion-6b · How It Works · Data refreshed daily, snapshot 2026-09-23.