calme-2.5-qwen2-7B — benchmark results

MaziyarPanahi's Qwen2 7B fine-tune from the Calme-2 series, a sibling of calme-2.1 aiming to improve the base model across benchmarks. Provider: Other. Released 2024-06-27. Access: Open.

Unified ELO 1431 ± 50, rank #1086 of 1776 rated models, from 7 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard - MuSR15.79Score84
Open LLM Leaderboard - MATH Level 522.58Score75.8
Open Korean LLM Leaderboard324.1Average Score (%)68.3
Open LLM Leaderboard - GPQA8.05Score67.9
Open LLM Leaderboard - MMLU-Pro29.8Score58.9
Open LLM Leaderboard - BBH28.28Score46.2
Open LLM Leaderboard - IFEval31.45Score29.1

Interactive version: theaggregate.ai/model?slug=calme-2-5-qwen2-7b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.