Claude Opus 4.1: benchmark results

Anthropic Opus-tier Claude model, an incremental Opus 4 upgrade focused on coding and agentic tasks. Provider: Anthropic. Released 2025-08-05. Access: API.

Unified ELO 1657 ± 1, rank #114 of 1392 rated models, from 144 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
FutureSearch DRB - Find Dataset0.71Average Score100
Nejumi 4 - ALT - Truthfulness88.5Score (%)100
PCB-Bench - Placement Macro CQ93.3Accuracy (%)100
PCB-Bench - Placement Micro CQ94.35Accuracy (%)100
PCB-Bench - Routing Micro CQ92.32Accuracy (%)100
PCB-Bench - Routing Micro QA SBERT57.38SBERT similarity (%)100
Nejumi 4 - ALT - Toxicity87.79Score (%)99.2
Phare - Jailbreak Resistance81.35Score (%)98.5
BenchTable81.9Total Score (%)98.3
Nejumi 4 - GLP - Function Calling71.42Score (%)97.5
SciArena1126.4Elo Rating97.3
PCB-Bench - Routing Macro CQ99.16Accuracy (%)95.8

Interactive version: theaggregate.ai/model?slug=claude-opus-4-1 · How It Works · Data refreshed daily, snapshot 2026-09-05.