c4ai-command-a-03-2025 — benchmark results
Cohere's open-weights 111B dense enterprise model (256K context) tuned for agents, RAG, and 23 languages; runs on two GPUs (March 2025). Provider: Cohere. Released 2025-03-13. Access: Open.
Unified ELO 1544 ± 21, rank #640 of 1776 rated models, from 29 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BenCzechMark | 82.29 | Average Score (%) | 95.5 |
| UGI - Natural Intelligence | 32.78 | NatInt Score | 79.6 |
| UGI - Writing | 40.59 | Writing Score | 75.6 |
| Arabic Broad Leaderboard | 8.72 | Average Score (0-10) | 72.4 |
| Hebrew LLM - Winograd (0-shot) | 84.89 | Accuracy (%) | 61.3 |
| K-MetBench | 65.5 | Accuracy (self-reported) | 60.3 |
| Mizan LLM Leaderboard | 61 | Average Score (0-100) | 52.3 |
| Hebrew LLM - Trivia (0-shot) | 63.79 | Accuracy (%) | 51.6 |
| Hebrew LLM - Nikud | 61.12 | Score (%) | 48.4 |
| RewardBench 2 Focus | 86.67 | Accuracy (%) | 41.2 |
| GSM-MC | 98.18 | Accuracy (%) | 38.8 |
| MATH-MC Level 1 | 96.28 | Accuracy (%) | 32.4 |
Interactive version: theaggregate.ai/model?slug=c4ai-command-a-03-2025 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.