Command R: benchmark results
Cohere Command R model optimized for retrieval-augmented generation and enterprise chat. Provider: Cohere. Released 2024-03-11. Access: Open.
Unified ELO 1416 ± 1, rank #1083 of 1392 rated models, from 74 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MERA - SimpleAr | 99.4 | EM (%) | 74 |
| MERA - ruDetox | 30.14 | Joint Score (%) | 73.1 |
| MERA - PARus | 89.2 | Accuracy (%) | 71.1 |
| MERA - MultiQ | 50.58 | F1 (%) | 68.9 |
| MERA - CheGeKa | 29.75 | F1 (%) | 67.5 |
| RULER | 88.9 | Avg Accuracy (%) | 61.3 |
| CanAiCode | 98.9 | Junior-v2 Python Pass Rate (%) | 60.2 |
| MERA - ruHateSpeech | 78.11 | Accuracy (%) | 60.1 |
| MERA - ruTiE | 74.96 | Accuracy (%) | 55 |
| HELM NaturalQuestions (Open) | 72.05 | F1 (%) | 53.3 |
| HELM NarrativeQA | 74.17 | F1 (%) | 50 |
| IFEval Leaderboard | 69.71 | Final Score | 50 |
Interactive version: theaggregate.ai/model?slug=command-r · How It Works · Data refreshed daily, snapshot 2026-09-05.