c4ai-command-a-03-2025 — benchmark results

Cohere's open-weights 111B dense enterprise model (256K context) tuned for agents, RAG, and 23 languages; runs on two GPUs (March 2025). Provider: Cohere. Released 2025-03-13. Access: Open.

Unified ELO 1544 ± 21, rank #640 of 1776 rated models, from 29 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BenCzechMark82.29Average Score (%)95.5
UGI - Natural Intelligence32.78NatInt Score79.6
UGI - Writing40.59Writing Score75.6
Arabic Broad Leaderboard8.72Average Score (0-10)72.4
Hebrew LLM - Winograd (0-shot)84.89Accuracy (%)61.3
K-MetBench65.5Accuracy (self-reported)60.3
Mizan LLM Leaderboard61Average Score (0-100)52.3
Hebrew LLM - Trivia (0-shot)63.79Accuracy (%)51.6
Hebrew LLM - Nikud61.12Score (%)48.4
RewardBench 2 Focus86.67Accuracy (%)41.2
GSM-MC98.18Accuracy (%)38.8
MATH-MC Level 196.28Accuracy (%)32.4

Interactive version: theaggregate.ai/model?slug=c4ai-command-a-03-2025 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.