command-a-reasoning-08-2025 — benchmark results

Cohere's open-weights Command A Reasoning release (August 2025). Provider: Cohere. Released 2025-08-21. Access: Open.

Unified ELO 1625 ± 18, rank #425 of 1776 rated models, from 16 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
K-MetBench77.8Accuracy (self-reported)87.9
RewardBench 2 Safety88.61Accuracy (%)78.4
RewardBench 2 Math86.89Accuracy (%)56.9
RewardBench 2 Focus88.94Accuracy (%)52.9
JudgeBench Math87.5Accuracy (%)46.1
JudgeBench Coding89.29Accuracy (%)44.1
RewardBench 2 Factuality65.63Accuracy (%)40.2
JudgeBench Reasoning84.69Accuracy (%)37.3
RewardBench 2 Precise IF45Accuracy (%)37.3
GSM-MC97.95Accuracy (%)36.6
JudgeBench Knowledge71.1Accuracy (%)35.3
MATH-MC Level 196.05Accuracy (%)30.1

Interactive version: theaggregate.ai/model?slug=command-a-reasoning-08-2025 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.