Command: benchmark results
Provider: Cohere. Released 2023-09-29. Access: API.
Unified ELO 1397 ± 1, rank #1150 of 1392 rated models, from 8 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM NaturalQuestions (Open) | 77.69 | F1 (%) | 88.9 |
| HELM NarrativeQA | 74.88 | F1 (%) | 60 |
| HELM NaturalQuestions (Closed) | 39.11 | F1 (%) | 60 |
| HELM Lite | 35.79 | Mean win rate (self-reported) | 30.3 |
| HELM (Stanford) | 32.68 | Mean Win Rate (%) | 26.7 |
| CanAiCode | 93.4 | Junior-v2 Python Pass Rate (%) | 24.5 |
| HELM WMT 2014 | 8.82 | BLEU-4 (%) | 4.4 |
| Conceptual Reasoning Index - Decision Theory (DTBench) | 0.8 | Chance-Corrected Score (0-100) | 0 |
Interactive version: theaggregate.ai/model?slug=command · How It Works · Data refreshed daily, snapshot 2026-09-05.