command-a-reasoning-08-2025 — benchmark results
Cohere's open-weights Command A Reasoning release (August 2025). Provider: Cohere. Released 2025-08-21. Access: Open.
Unified ELO 1625 ± 18, rank #425 of 1776 rated models, from 16 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| K-MetBench | 77.8 | Accuracy (self-reported) | 87.9 |
| RewardBench 2 Safety | 88.61 | Accuracy (%) | 78.4 |
| RewardBench 2 Math | 86.89 | Accuracy (%) | 56.9 |
| RewardBench 2 Focus | 88.94 | Accuracy (%) | 52.9 |
| JudgeBench Math | 87.5 | Accuracy (%) | 46.1 |
| JudgeBench Coding | 89.29 | Accuracy (%) | 44.1 |
| RewardBench 2 Factuality | 65.63 | Accuracy (%) | 40.2 |
| JudgeBench Reasoning | 84.69 | Accuracy (%) | 37.3 |
| RewardBench 2 Precise IF | 45 | Accuracy (%) | 37.3 |
| GSM-MC | 97.95 | Accuracy (%) | 36.6 |
| JudgeBench Knowledge | 71.1 | Accuracy (%) | 35.3 |
| MATH-MC Level 1 | 96.05 | Accuracy (%) | 30.1 |
Interactive version: theaggregate.ai/model?slug=command-a-reasoning-08-2025 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.