japanese-stablelm-instruct-gamma-7B: benchmark results
Provider: Other. Access: Open.
Unified ELO 1406 ± 20, rank #2088 of 2928 rated models, from 10 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| pfgen-bench - QA Mode - Helpfulness | 0.16 | Helpfulness Score | 46.1 |
| pfgen-bench - QA Mode - Truthfulness | 0.73 | Truthfulness Score | 45.8 |
| pfgen-bench - QA Mode - Score | 0.48 | pfgen Score (mean of three) | 44.8 |
| pfgen-bench - QA Mode - Fluency | 0.55 | Fluency Score | 43.6 |
| Open LLM Leaderboard v1 - GSM8K | 19.26 | Accuracy (%) (5-shot) | 40.5 |
| Open LLM Leaderboard v1 - HellaSwag | 78.68 | Normalized accuracy (%) (10-shot) | 35.3 |
| Open LLM Leaderboard v1 - MMLU | 54.82 | Accuracy (%) (5-shot) | 34.8 |
| Open LLM Leaderboard v1 - WinoGrande | 73.72 | Accuracy (%) (5-shot) | 33.3 |
| Open LLM Leaderboard v1 - ARC Challenge | 50.68 | Normalized accuracy (%) (25-shot) | 25 |
| Open LLM Leaderboard v1 - TruthfulQA MC2 | 39.77 | MC2 (%) (0-shot) | 11.9 |
Interactive version: theaggregate.ai/model?slug=japanese-stablelm-instruct-gamma-7b · How It Works · Data refreshed daily, snapshot 2026-09-23.