Command A 03 2025: benchmark results

Provider: Cohere. Released 2025-03-13. Access: Open.

Unified ELO 1631 ± 35, rank #517 of 2656 rated models, from 31 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BALSAM - Factuality26.56Overall score (0-100, LLM-judged generation and multiple cho92.9
BALSAM - Reading Comprehension59.8Overall score (0-100, LLM-judged generation and multiple cho89.3
BenchTable - Tech79.7Weighted Score (%)87.5
BALSAM - Text Manipulation64.34Overall score (0-100, LLM-judged generation and multiple cho85.7
BALSAM - Text Classification39.82Overall score (0-100, LLM-judged generation and multiple cho82.1
Vals AI CaseLaw v264.52Accuracy (%)78.8
BALSAM - Creative Writing51.1Overall score (0-100, LLM-judged generation and multiple cho78.6
BALSAM - Translation/Transliteration64.95Overall score (0-100, LLM-judged generation and multiple cho78.6
BALSAM - Question Answering72.2Overall score (0-100, LLM-judged generation and multiple cho75
BALSAM - Overall55.16Mean of category overall scores (0-100)70.4
BenchTable - Utility62.8Weighted Score (%)68.9
BALSAM - Information Extraction49.92Overall score (0-100, LLM-judged generation and multiple cho67.9

Interactive version: theaggregate.ai/model?slug=command-a-03-2025 · How It Works · Data refreshed daily, snapshot 2026-09-19.