calm2-7B-chat-dpo-experimental: benchmark results

Provider: Other. Access: Open.

Unified ELO 1323 ± 20, rank #2605 of 2928 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
pfgen-bench - Completion Mode - Score0.63pfgen Score (mean of three)78.9
pfgen-bench - Completion Mode - Fluency0.74Fluency Score78.6
pfgen-bench - Completion Mode - Truthfulness0.86Truthfulness Score78.4
pfgen-bench - Completion Mode - Helpfulness0.29Helpfulness Score78.3
pfgen-bench - QA Mode - Fluency0.64Fluency Score64.5
pfgen-bench - QA Mode - Helpfulness0.24Helpfulness Score61.8
pfgen-bench - QA Mode - Score0.55pfgen Score (mean of three)61.2
pfgen-bench - QA Mode - Truthfulness0.78Truthfulness Score52.7
Open LLM Leaderboard v1 - GSM8K5.53Accuracy (%) (5-shot)24.4
Open LLM Leaderboard v1 - TruthfulQA MC243.13MC2 (%) (0-shot)22.6
Open LLM Leaderboard v1 - MMLU39.82Accuracy (%) (5-shot)21
Open LLM Leaderboard v1 - HellaSwag68.99Normalized accuracy (%) (10-shot)19.8

Interactive version: theaggregate.ai/model?slug=calm2-7b-chat-dpo-experimental · How It Works · Data refreshed daily, snapshot 2026-09-23.