Swallow-70B-instruct-hf: benchmark results
Provider: Other. Access: Open.
Unified ELO 1492 ± 20, rank #1339 of 2928 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| pfgen-bench - Completion Mode - Fluency | 0.87 | Fluency Score | 98.1 |
| pfgen-bench - Completion Mode - Score | 0.77 | pfgen Score (mean of three) | 97.8 |
| pfgen-bench - Completion Mode - Helpfulness | 0.49 | Helpfulness Score | 97.4 |
| pfgen-bench - Completion Mode - Truthfulness | 0.93 | Truthfulness Score | 96 |
| Open LLM Leaderboard v1 - MMLU | 67.08 | Accuracy (%) (5-shot) | 88.8 |
| Open LLM Leaderboard v1 - WinoGrande | 82.08 | Accuracy (%) (5-shot) | 82.2 |
| Open LLM Leaderboard v1 - ARC Challenge | 66.21 | Normalized accuracy (%) (25-shot) | 72.7 |
| Open LLM Leaderboard v1 - HellaSwag | 85.14 | Normalized accuracy (%) (10-shot) | 72.1 |
| pfgen-bench - QA Mode - Helpfulness | 0.33 | Helpfulness Score | 71.5 |
| pfgen-bench - QA Mode - Score | 0.57 | pfgen Score (mean of three) | 64.8 |
| pfgen-bench - QA Mode - Fluency | 0.64 | Fluency Score | 63 |
| Open LLM Leaderboard v1 - GSM8K | 45.94 | Accuracy (%) (5-shot) | 59.8 |
Interactive version: theaggregate.ai/model?slug=swallow-70b-instruct-hf · How It Works · Data refreshed daily, snapshot 2026-09-23.