Command A 03 2025: benchmark results
Provider: Cohere. Released 2025-03-13. Access: Open.
Unified ELO 1631 ± 35, rank #517 of 2656 rated models, from 31 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BALSAM - Factuality | 26.56 | Overall score (0-100, LLM-judged generation and multiple cho | 92.9 |
| BALSAM - Reading Comprehension | 59.8 | Overall score (0-100, LLM-judged generation and multiple cho | 89.3 |
| BenchTable - Tech | 79.7 | Weighted Score (%) | 87.5 |
| BALSAM - Text Manipulation | 64.34 | Overall score (0-100, LLM-judged generation and multiple cho | 85.7 |
| BALSAM - Text Classification | 39.82 | Overall score (0-100, LLM-judged generation and multiple cho | 82.1 |
| Vals AI CaseLaw v2 | 64.52 | Accuracy (%) | 78.8 |
| BALSAM - Creative Writing | 51.1 | Overall score (0-100, LLM-judged generation and multiple cho | 78.6 |
| BALSAM - Translation/Transliteration | 64.95 | Overall score (0-100, LLM-judged generation and multiple cho | 78.6 |
| BALSAM - Question Answering | 72.2 | Overall score (0-100, LLM-judged generation and multiple cho | 75 |
| BALSAM - Overall | 55.16 | Mean of category overall scores (0-100) | 70.4 |
| BenchTable - Utility | 62.8 | Weighted Score (%) | 68.9 |
| BALSAM - Information Extraction | 49.92 | Overall score (0-100, LLM-judged generation and multiple cho | 67.9 |
Interactive version: theaggregate.ai/model?slug=command-a-03-2025 · How It Works · Data refreshed daily, snapshot 2026-09-19.