Claude Opus 4 (Thinking 16K) — benchmark results

Claude Opus 4 evaluated with a 16K-token thinking budget. Provider: Anthropic. Released 2025-05-22. Access: API.

Unified ELO 1723 ± 34, rank #235 of 1776 rated models, from 12 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Generalization V1 (Lechmazur)1.69Avg Rank (lower is better)98.8
Confabulation Leaderboard (Lechmazur)2.48Confabulation rate % (lower is better)98.4
Chatbot Arena (Text)1424Elo75.4
NYT Connections Older Models49.7Score (%)73.1
ARC-AGI-28.61Accuracy (%)57.7
Chatbot Arena (Vision)1206Arena Score56.3
Step Game (Lechmazur)2.4TrueSkill μ52.7
WeirdML43.72Average Score51.1
OTIS Mock AIME 2024-2560Accuracy (%)48.7
Elimination Game (Lechmazur)3.86TrueSkill μ45.8
ARC-AGI-135.67Accuracy (%)37.3
VPCT38Accuracy (%)35.9

Interactive version: theaggregate.ai/model?slug=claude-opus-4-thinking-16k · How the rankings work · Data refreshed daily, snapshot 2026-07-22.