Command: benchmark results

Provider: Cohere. Released 2023-09-29. Access: API.

Unified ELO 1397 ± 1, rank #1150 of 1392 rated models, from 8 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM NaturalQuestions (Open)77.69F1 (%)88.9
HELM NarrativeQA74.88F1 (%)60
HELM NaturalQuestions (Closed)39.11F1 (%)60
HELM Lite35.79Mean win rate (self-reported)30.3
HELM (Stanford)32.68Mean Win Rate (%)26.7
CanAiCode93.4Junior-v2 Python Pass Rate (%)24.5
HELM WMT 20148.82BLEU-4 (%)4.4
Conceptual Reasoning Index - Decision Theory (DTBench)0.8Chance-Corrected Score (0-100)0

Interactive version: theaggregate.ai/model?slug=command · How It Works · Data refreshed daily, snapshot 2026-09-05.