Claude Opus 4 (2025-05-14) (Thinking): benchmark results

Provider: Anthropic. Released 2025-05-22. Access: API.

Unified ELO 1637 ± 1, rank #310 of 1919 rated models, from 12 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Capabilities - MMLU-Pro87.5COT correct100
HELM Capabilities - Omni-MATH61.6Acc94
HELM Capabilities - GPQA70.85COT correct92
HELM Capabilities - WildBench85.19WB Score90
HELM Safety - Bbq97.1BBQ accuracy88.9
HELM Safety96.94Mean score (self-reported)87
HELM Safety - Safety Scenarios96.94Mean score86
HELM Safety - Simple Safety Tests100LM Evaluated Safety score84.6
HELM Safety - Harm Bench91.69LM Evaluated Safety score75.5
HELM Capabilities - IFEval84.87IFEval Strict Acc74
HELM Safety - Xstest97LM Evaluated Safety score69.2
HELM Safety - Anthropic Red Team98.9LM Evaluated Safety score49.5

Interactive version: theaggregate.ai/model?slug=claude-opus-4-2025-05-14-thinking · How It Works · Data refreshed daily, snapshot 2026-09-08.