calm2-7B-chat-dpo-experimental: benchmark results
Provider: Other. Access: Open.
Unified ELO 1323 ± 20, rank #2605 of 2928 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| pfgen-bench - Completion Mode - Score | 0.63 | pfgen Score (mean of three) | 78.9 |
| pfgen-bench - Completion Mode - Fluency | 0.74 | Fluency Score | 78.6 |
| pfgen-bench - Completion Mode - Truthfulness | 0.86 | Truthfulness Score | 78.4 |
| pfgen-bench - Completion Mode - Helpfulness | 0.29 | Helpfulness Score | 78.3 |
| pfgen-bench - QA Mode - Fluency | 0.64 | Fluency Score | 64.5 |
| pfgen-bench - QA Mode - Helpfulness | 0.24 | Helpfulness Score | 61.8 |
| pfgen-bench - QA Mode - Score | 0.55 | pfgen Score (mean of three) | 61.2 |
| pfgen-bench - QA Mode - Truthfulness | 0.78 | Truthfulness Score | 52.7 |
| Open LLM Leaderboard v1 - GSM8K | 5.53 | Accuracy (%) (5-shot) | 24.4 |
| Open LLM Leaderboard v1 - TruthfulQA MC2 | 43.13 | MC2 (%) (0-shot) | 22.6 |
| Open LLM Leaderboard v1 - MMLU | 39.82 | Accuracy (%) (5-shot) | 21 |
| Open LLM Leaderboard v1 - HellaSwag | 68.99 | Normalized accuracy (%) (10-shot) | 19.8 |
Interactive version: theaggregate.ai/model?slug=calm2-7b-chat-dpo-experimental · How It Works · Data refreshed daily, snapshot 2026-09-23.