sarashina2.2-3B: benchmark results
Provider: Other. Access: Open.
Unified ELO 1533 ± 44, rank #926 of 2656 rated models, from 26 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| pfgen-bench - Completion Mode - Helpfulness | 0.47 | Helpfulness Score | 95.2 |
| pfgen-bench - Completion Mode - Score | 0.73 | pfgen Score (mean of three) | 93.3 |
| pfgen-bench - Completion Mode - Fluency | 0.82 | Fluency Score | 91.1 |
| pfgen-bench - Completion Mode - Truthfulness | 0.91 | Truthfulness Score | 89.6 |
| Swallow - Pre-trained Japanese - NIILC | 64.2 | Character F1 (%) | 79.6 |
| Swallow - Pre-trained English - HumanEval | 53 | Pass@1 (%) | 75.5 |
| Swallow - Pre-trained Japanese - JEMHopQA | 56.3 | Character F1 (%) | 72.4 |
| Swallow - Pre-trained Japanese - MGSM | 59.6 | Exact Match (%) | 70.4 |
| Swallow - Pre-trained Japanese - WMT20 En-Ja | 27.3 | BLEU | 67.3 |
| Swallow - Pre-trained Japanese - Average | 51.6 | Average Score (%) | 65.3 |
| Swallow - Pre-trained Japanese - JCommonsenseQA | 91.1 | Accuracy (%) | 65.3 |
| Swallow - Pre-trained Japanese - JSQuAD | 90.6 | Character F1 (%) | 59.2 |
Interactive version: theaggregate.ai/model?slug=sarashina2-2-3b · How It Works · Data refreshed daily, snapshot 2026-09-19.