Claude Opus 4.7 (Thinking): benchmark results
Claude Opus 4.7 evaluated with thinking enabled. Provider: Anthropic. Released 2026-04-16. Access: API.
Unified ELO 1727 ± 1, rank #57 of 1761 rated models, from 23 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Chinese Classical Bench - Char-Gloss Judge | 73.6 | Score (%) | 100 |
| Chinese Classical Bench - Fill-In Exact | 88 | Score (%) | 100 |
| Chinese Classical Bench - Translate Judge | 80.2 | Score (%) | 100 |
| KnotBench | 54.6 | Accuracy (%) (self-reported) | 100 |
| Wolfram LLM Benchmarking Project | 72.5 | Correct Functionality (%) | 99.8 |
| ProfBench | 61.3 | Overall Rubric Score (%) | 99 |
| Kagi LLM Benchmark | 80.7 | Accuracy (%) | 95.8 |
| Design Arena (ASCII Art) | 1307 | Elo | 95.1 |
| ALE-Bench | 1323.05 | Performance (Self-Refine x1) (self-reported) | 93.2 |
| Design Arena (Data Viz) | 1306 | Elo | 92.9 |
| Design Arena (Website) | 1309 | Elo | 92.5 |
| Design Arena (UI Components) | 1329 | Elo | 92.2 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-7-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.