calme-3.1-instruct-3B — benchmark results
MaziyarPanahi's general-purpose fine-tune of Qwen2.5 3B, trained on EvolKit and French instruction data. Provider: Other. Released 2024-11-07. Access: Open.
Unified ELO 1404 ± 19, rank #1206 of 1776 rated models, from 10 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - MATH Level 5 | 17.75 | Score | 66.9 |
| French LLM Leaderboard - BAC FR | 30.08 | Score (%) | 59.3 |
| Open LLM Leaderboard - MMLU-Pro | 28.41 | Score | 53.3 |
| Open LLM Leaderboard - IFEval | 43.36 | Score | 45.8 |
| Open LLM Leaderboard - BBH | 27.31 | Score | 43.4 |
| Open LLM Leaderboard - GPQA | 4.81 | Score | 41 |
| French LLM Leaderboard - IFEval FR | 38.78 | Score (%) | 40.7 |
| French LLM Leaderboard - Average | 31.2 | Average Score (%) | 38.9 |
| Open LLM Leaderboard - MuSR | 7.4 | Score | 34.6 |
| French LLM Leaderboard - GPQA FR | 24.73 | Score (%) | 11.1 |
Interactive version: theaggregate.ai/model?slug=calme-3-1-instruct-3b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.