c4ai-command-r7B-12-2024 — benchmark results
Cohere's smallest Command R (7B) with a 128K context, open weights tuned for RAG, tool use, and agents. Provider: Cohere. Released 2024-12-11. Access: Open.
Unified ELO 1436 ± 22, rank #1065 of 1776 rated models, from 51 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - IFEval | 77.13 | Score | 93.1 |
| Open LLM Leaderboard - MATH Level 5 | 29.91 | Score | 81.8 |
| EuroEval Spanish NLU - Sentiment Headlines ES | 47.95 | Sentiment classification Score (%) | 80.8 |
| Arabic IFEval | 60.8 | Arabic Accuracy (%) | 79.5 |
| EuroEval Spanish NLU | 50.76 | NLU Average Score (%) | 77.3 |
| EuroEval Spanish NLU - MLQA ES | 64.05 | Reading comprehension Score (%) | 75.9 |
| Open LLM Leaderboard - BBH | 36.02 | Score | 73.5 |
| EuroEval Portuguese NLU - HAREM | 48.11 | Named entity recognition Score (%) | 71.7 |
| EuroEval Italian NLU - MultiNERD IT | 69.8 | Named entity recognition Score (%) | 71.2 |
| EuroEval Portuguese NLU - ScaLA PT | 20.49 | Linguistic acceptability Score (%) | 70.9 |
| EuroEval Italian NLU - Sentipolc16 | 58.28 | Sentiment classification Score (%) | 70.4 |
| EuroEval Italian NLU | 54 | NLU Average Score (%) | 67.2 |
Interactive version: theaggregate.ai/model?slug=c4ai-command-r7b-12-2024 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.