sarashina2.2-0.5B: benchmark results
Provider: Other. Access: Open.
Unified ELO 1346 ± 47, rank #2257 of 2656 rated models, from 26 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| pfgen-bench - Completion Mode - Fluency | 0.61 | Fluency Score | 59.6 |
| pfgen-bench - Completion Mode - Score | 0.5 | pfgen Score (mean of three) | 56.5 |
| pfgen-bench - Completion Mode - Truthfulness | 0.76 | Truthfulness Score | 54.3 |
| pfgen-bench - Completion Mode - Helpfulness | 0.12 | Helpfulness Score | 51.4 |
| Swallow - Pre-trained Japanese - JEMHopQA | 47.2 | Character F1 (%) | 41.8 |
| Swallow - Pre-trained Japanese - NIILC | 45.1 | Character F1 (%) | 38.8 |
| Swallow - Pre-trained Japanese - MGSM | 19.6 | Exact Match (%) | 32.7 |
| Swallow - Pre-trained Japanese - WMT20 En-Ja | 20.1 | BLEU | 27.6 |
| Swallow - Pre-trained Japanese - JSQuAD | 82.4 | Character F1 (%) | 24.5 |
| Swallow - Pre-trained English - GSM8K | 24.6 | Exact Match (%) | 22.4 |
| Swallow - Pre-trained English - HumanEval | 22.3 | Pass@1 (%) | 22.4 |
| Swallow - Pre-trained English - MATH | 13 | Exact Match (%) | 20.4 |
Interactive version: theaggregate.ai/model?slug=sarashina2-2-0-5b · How It Works · Data refreshed daily, snapshot 2026-09-19.