c4ai-command-a-03-2025: benchmark results

Cohere's open-weights 111B dense enterprise model (256K context) tuned for agents, RAG, and 23 languages; runs on two GPUs (March 2025). Provider: Cohere. Released 2025-03-13. Access: Open.

Unified ELO 1564 ± 1, rank #387 of 1392 rated models, from 31 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BenCzechMark78.09Average Score (%)94.5
UGI - Natural Intelligence32.78NatInt Score78.8
UGI - Writing40.59Writing Score74.9
Arabic Broad Leaderboard8.72Average Score (0-10)71.6
Hebrew LLM - Winograd (0-shot)84.89Accuracy (%)60.6
K-MetBench65.5Accuracy (self-reported)60.3
Mizan LLM Leaderboard61Average Score (0-100)52.3
Hebrew LLM - Nikud61.12Score (%)51.5
Hebrew LLM - Trivia (0-shot)63.79Accuracy (%)51.5
RewardBench 2 Focus86.67Accuracy (%)41.2
GSM-MC98.18Accuracy (%)38.8
MATH-MC Level 196.28Accuracy (%)32.4

Interactive version: theaggregate.ai/model?slug=c4ai-command-a-03-2025 · How It Works · Data refreshed daily, snapshot 2026-09-05.