c4ai-command-a-03-2025: benchmark results
Cohere's open-weights 111B dense enterprise model (256K context) tuned for agents, RAG, and 23 languages; runs on two GPUs (March 2025). Provider: Cohere. Released 2025-03-13. Access: Open.
Unified ELO 1564 ± 1, rank #387 of 1392 rated models, from 31 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BenCzechMark | 78.09 | Average Score (%) | 94.5 |
| UGI - Natural Intelligence | 32.78 | NatInt Score | 78.8 |
| UGI - Writing | 40.59 | Writing Score | 74.9 |
| Arabic Broad Leaderboard | 8.72 | Average Score (0-10) | 71.6 |
| Hebrew LLM - Winograd (0-shot) | 84.89 | Accuracy (%) | 60.6 |
| K-MetBench | 65.5 | Accuracy (self-reported) | 60.3 |
| Mizan LLM Leaderboard | 61 | Average Score (0-100) | 52.3 |
| Hebrew LLM - Nikud | 61.12 | Score (%) | 51.5 |
| Hebrew LLM - Trivia (0-shot) | 63.79 | Accuracy (%) | 51.5 |
| RewardBench 2 Focus | 86.67 | Accuracy (%) | 41.2 |
| GSM-MC | 98.18 | Accuracy (%) | 38.8 |
| MATH-MC Level 1 | 96.28 | Accuracy (%) | 32.4 |
Interactive version: theaggregate.ai/model?slug=c4ai-command-a-03-2025 · How It Works · Data refreshed daily, snapshot 2026-09-05.