Command-R+: benchmark results
Cohere Command R+ model, a higher-capacity Command R variant for retrieval, tool use, and complex enterprise tasks. Provider: Cohere. Released 2024-04-04. Access: Open.
Unified ELO 1458 ± 1, rank #906 of 1392 rated models, from 79 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Safety SimpleSafetyTests | 100 | LM Evaluated Safety score (%) | 86 |
| CRM LLM Leaderboard | 66.2 | CRM Accuracy (0-100) | 85.3 |
| RABBITS B4BQA | 98.28 | Accuracy (%) | 73.2 |
| Ru Arena Hard | 77.17 | Win Rate (%) | 71.9 |
| IFEval Leaderboard | 77.01 | Final Score | 70 |
| BenchBench | 61.83 | Aggregate Score (%) | 66.9 |
| MixEval | 51.4 | Score | 64.7 |
| HELM WMT 2014 | 20.33 | BLEU-4 (%) | 62.8 |
| NoLiMa | 90.9 | Base Score (%) | 61.9 |
| MAGI-Hard | 49.7 | Accuracy (%, 2024-05 snapshot) | 60.3 |
| LingOly | 21.5 | Exact Match Accuracy | 60 |
| EvalPlus (HumanEval+ & MBPP+) | 60.1 | Pass@1 avg (%) | 58.1 |
Interactive version: theaggregate.ai/model?slug=command-r-plus · How It Works · Data refreshed daily, snapshot 2026-09-05.