Claude Opus 4.8 (Thinking) — benchmark results

Claude Opus 4.8 evaluated with thinking enabled. Provider: Anthropic. Released 2026-05-29. Access: API.

Unified ELO 1916 ± 26, rank #53 of 1776 rated models, from 17 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Kagi LLM Benchmark88.8Accuracy (%)98.6
Wolfram LLM Benchmarking Project68.1Correct Functionality (%)97.2
Chatbot Arena (Text)1484Elo96.8
Chatbot Arena (Code)1565Elo96
Agent Arena - Tool Hallucination0.22Tool Hallucination (%)94.6
GeoBench Photos4101.62Average Score93.8
Chatbot Arena (Vision)1286Arena Score93.3
ProfBench56.8Overall Rubric Score (%)92
Agent Arena - Praise vs Complaint19.42Praise vs Complaint (%)91.9
GeoBench ACW4127.75Average Score87.2
PM-LLM-Benchmark35.3Score86.2
ProphetArena0.951 - Brier Score85.4

Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.