Claude Opus 4.1 — benchmark results

Anthropic Opus-tier Claude model, an incremental Opus 4 upgrade focused on coding and agentic tasks. Provider: Anthropic. Released 2025-08-05. Access: API.

Unified ELO 1698 ± 11, rank #270 of 1776 rated models, from 123 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
FutureSearch DRB - Find Dataset0.71Average Score100
PCB-Bench - Placement Macro CQ93.3Accuracy (%)100
PCB-Bench - Placement Micro CQ94.35Accuracy (%)100
PCB-Bench - Routing Micro CQ92.32Accuracy (%)100
PCB-Bench - Routing Micro QA SBERT57.38SBERT similarity (%)100
BenchTable81.9Total Score (%)98.3
SciArena1126.4Elo Rating97.3
Vals AI MGSM94.44Accuracy (%)97
PCB-Bench - Routing Macro CQ99.16Accuracy (%)95.8
GSMA Open-Telco - ORAN-Bench87.33Score (%)95.3
Vals AI MATH 50095.4Accuracy (%)94.1
LingOly-TOO45.8Obfuscated Score (self-reported)93.3

Interactive version: theaggregate.ai/model?slug=claude-opus-4-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.