sarashina2-70B: benchmark results
Provider: Other. Access: Open.
Unified ELO 1582 ± 45, rank #702 of 2656 rated models, from 26 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Swallow - Pre-trained Japanese - JEMHopQA | 71.7 | Character F1 (%) | 100 |
| pfgen-bench - Completion Mode - Helpfulness | 0.49 | Helpfulness Score | 96.8 |
| Swallow - Pre-trained English - SQuAD2 | 67.5 | Exact Match (%) | 95.9 |
| Swallow - Pre-trained Japanese - JSQuAD | 92.9 | Character F1 (%) | 95.9 |
| pfgen-bench - Completion Mode - Score | 0.74 | pfgen Score (mean of three) | 95.2 |
| Swallow - Pre-trained Japanese - WMT20 En-Ja | 31.3 | BLEU | 94.9 |
| pfgen-bench - Completion Mode - Truthfulness | 0.92 | Truthfulness Score | 94.7 |
| pfgen-bench - Completion Mode - Fluency | 0.83 | Fluency Score | 93.3 |
| Swallow - Pre-trained Japanese - NIILC | 66.8 | Character F1 (%) | 91.8 |
| Swallow - Pre-trained English - XWINO | 91.7 | Accuracy (%) | 90.8 |
| Swallow - Pre-trained Japanese - WMT20 Ja-En | 24.3 | BLEU | 81.6 |
| Swallow - Pre-trained English - HellaSwag | 62.8 | Accuracy (%) | 76.5 |
Interactive version: theaggregate.ai/model?slug=sarashina2-70b · How It Works · Data refreshed daily, snapshot 2026-09-19.