Claude Opus 4.1 — benchmark results
Anthropic Opus-tier Claude model, an incremental Opus 4 upgrade focused on coding and agentic tasks. Provider: Anthropic. Released 2025-08-05. Access: API.
Unified ELO 1698 ± 11, rank #270 of 1776 rated models, from 123 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| FutureSearch DRB - Find Dataset | 0.71 | Average Score | 100 |
| PCB-Bench - Placement Macro CQ | 93.3 | Accuracy (%) | 100 |
| PCB-Bench - Placement Micro CQ | 94.35 | Accuracy (%) | 100 |
| PCB-Bench - Routing Micro CQ | 92.32 | Accuracy (%) | 100 |
| PCB-Bench - Routing Micro QA SBERT | 57.38 | SBERT similarity (%) | 100 |
| BenchTable | 81.9 | Total Score (%) | 98.3 |
| SciArena | 1126.4 | Elo Rating | 97.3 |
| Vals AI MGSM | 94.44 | Accuracy (%) | 97 |
| PCB-Bench - Routing Macro CQ | 99.16 | Accuracy (%) | 95.8 |
| GSMA Open-Telco - ORAN-Bench | 87.33 | Score (%) | 95.3 |
| Vals AI MATH 500 | 95.4 | Accuracy (%) | 94.1 |
| LingOly-TOO | 45.8 | Obfuscated Score (self-reported) | 93.3 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.