sarashina2.2-3B: benchmark results

Provider: Other. Access: Open.

Unified ELO 1533 ± 44, rank #926 of 2656 rated models, from 26 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
pfgen-bench - Completion Mode - Helpfulness0.47Helpfulness Score95.2
pfgen-bench - Completion Mode - Score0.73pfgen Score (mean of three)93.3
pfgen-bench - Completion Mode - Fluency0.82Fluency Score91.1
pfgen-bench - Completion Mode - Truthfulness0.91Truthfulness Score89.6
Swallow - Pre-trained Japanese - NIILC64.2Character F1 (%)79.6
Swallow - Pre-trained English - HumanEval53Pass@1 (%)75.5
Swallow - Pre-trained Japanese - JEMHopQA56.3Character F1 (%)72.4
Swallow - Pre-trained Japanese - MGSM59.6Exact Match (%)70.4
Swallow - Pre-trained Japanese - WMT20 En-Ja27.3BLEU67.3
Swallow - Pre-trained Japanese - Average51.6Average Score (%)65.3
Swallow - Pre-trained Japanese - JCommonsenseQA91.1Accuracy (%)65.3
Swallow - Pre-trained Japanese - JSQuAD90.6Character F1 (%)59.2

Interactive version: theaggregate.ai/model?slug=sarashina2-2-3b · How It Works · Data refreshed daily, snapshot 2026-09-19.