sarashina2.2-3B-instruct-v0.1: benchmark results
Provider: Other. Access: Open.
Unified ELO 1504 ± 33, rank #1066 of 2656 rated models, from 47 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| pfgen-bench - Completion Mode - Helpfulness | 0.52 | Helpfulness Score | 98.7 |
| pfgen-bench - Completion Mode - Score | 0.75 | pfgen Score (mean of three) | 96.2 |
| pfgen-bench - Completion Mode - Fluency | 0.84 | Fluency Score | 94.9 |
| pfgen-bench - QA Mode - Fluency | 0.81 | Fluency Score | 89.1 |
| pfgen-bench - QA Mode - Helpfulness | 0.46 | Helpfulness Score | 87.3 |
| pfgen-bench - Completion Mode - Truthfulness | 0.89 | Truthfulness Score | 85.6 |
| pfgen-bench - Chat Mode - Fluency | 0.72 | Fluency Score | 85.3 |
| pfgen-bench - QA Mode - Score | 0.72 | pfgen Score (mean of three) | 84.8 |
| pfgen-bench - Chat Mode - Helpfulness | 0.3 | Helpfulness Score | 82.4 |
| pfgen-bench - Chat Mode - Score | 0.62 | pfgen Score (mean of three) | 79.4 |
| pfgen-bench - QA Mode - Truthfulness | 0.88 | Truthfulness Score | 73.9 |
| pfgen-bench - Chat Mode - Truthfulness | 0.83 | Truthfulness Score | 70.6 |
Interactive version: theaggregate.ai/model?slug=sarashina2-2-3b-instruct-v0-1 · How It Works · Data refreshed daily, snapshot 2026-09-19.