Claude Opus 4.7 (Thinking) — benchmark results
Claude Opus 4.7 evaluated with thinking enabled. Provider: Anthropic. Released 2026-04-16. Access: API.
Unified ELO 1889 ± 25, rank #70 of 1776 rated models, from 29 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Chinese Classical Bench - Char-Gloss Judge | 73.6 | Score (%) | 100 |
| Chinese Classical Bench - Fill-In Exact | 88 | Score (%) | 100 |
| Chinese Classical Bench - Translate Judge | 80.2 | Score (%) | 100 |
| ProfBench | 61.3 | Overall Rubric Score (%) | 100 |
| Wolfram LLM Benchmarking Project | 72.5 | Correct Functionality (%) | 99.8 |
| Chatbot Arena (Text) | 1502 | Elo | 99.5 |
| Chatbot Arena (Vision) | 1306 | Arena Score | 99.3 |
| Design Arena (UI Components) | 1345 | Elo | 97.8 |
| Design Arena (Game Dev) | 1336 | Elo | 96.4 |
| Kagi LLM Benchmark | 80.7 | Accuracy (%) | 95.7 |
| Design Arena (Website) | 1322 | Elo | 95.5 |
| Chatbot Arena (Code) | 1559 | Elo | 94.9 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-7-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.