sarashina2.2-1B: benchmark results

Provider: Other. Access: Open.

Unified ELO 1414 ± 44, rank #1797 of 2656 rated models, from 26 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
pfgen-bench - Completion Mode - Fluency0.69Fluency Score72.2
pfgen-bench - Completion Mode - Truthfulness0.83Truthfulness Score70.3
pfgen-bench - Completion Mode - Score0.58pfgen Score (mean of three)69.6
pfgen-bench - Completion Mode - Helpfulness0.22Helpfulness Score67.6
Swallow - Pre-trained Japanese - MGSM38.8Exact Match (%)49
Swallow - Pre-trained Japanese - NIILC52.3Character F1 (%)46.9
Swallow - Pre-trained English - HumanEval34.2Pass@1 (%)44.9
Swallow - Pre-trained English - GSM8K40.3Exact Match (%)38.8
Swallow - Pre-trained English - MATH20.6Exact Match (%)37.8
Swallow - Pre-trained Japanese - Average39.2Average Score (%)34.7
Swallow - Pre-trained Japanese - JEMHopQA46.2Character F1 (%)34.7
Swallow - Pre-trained Japanese - WMT20 En-Ja21.9BLEU34.7

Interactive version: theaggregate.ai/model?slug=sarashina2-2-1b · How It Works · Data refreshed daily, snapshot 2026-09-19.