Swallow-70B-instruct-hf: benchmark results

Provider: Other. Access: Open.

Unified ELO 1492 ± 20, rank #1339 of 2928 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
pfgen-bench - Completion Mode - Fluency0.87Fluency Score98.1
pfgen-bench - Completion Mode - Score0.77pfgen Score (mean of three)97.8
pfgen-bench - Completion Mode - Helpfulness0.49Helpfulness Score97.4
pfgen-bench - Completion Mode - Truthfulness0.93Truthfulness Score96
Open LLM Leaderboard v1 - MMLU67.08Accuracy (%) (5-shot)88.8
Open LLM Leaderboard v1 - WinoGrande82.08Accuracy (%) (5-shot)82.2
Open LLM Leaderboard v1 - ARC Challenge66.21Normalized accuracy (%) (25-shot)72.7
Open LLM Leaderboard v1 - HellaSwag85.14Normalized accuracy (%) (10-shot)72.1
pfgen-bench - QA Mode - Helpfulness0.33Helpfulness Score71.5
pfgen-bench - QA Mode - Score0.57pfgen Score (mean of three)64.8
pfgen-bench - QA Mode - Fluency0.64Fluency Score63
Open LLM Leaderboard v1 - GSM8K45.94Accuracy (%) (5-shot)59.8

Interactive version: theaggregate.ai/model?slug=swallow-70b-instruct-hf · How It Works · Data refreshed daily, snapshot 2026-09-23.