Claude Opus 5.5 (Low): benchmark results

Provider: Anthropic. Released 2026-09-22. Access: API.

Unified ELO 1884 ± 28, rank #46 of 2055 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EnigmaForge - Fact F191.9World-fact recovery F1 (%) over 600 instances, all items95.2
DuelLab Overall59.4DuelLab Score90
DataBench54Score (%)85.7
ARC-AGI-270.14Accuracy (%)80.3
ARC-AGI-188.5Accuracy (%)72.5
Bug Hunt Bench - VS Code Extension11.3Planted Bugs Fixed (out of 45)63.7
EnigmaForge - Intuition22Task success (%) on implicit-condition instances, where the 61.9
NonoBench60Overall Accuracy (%)60.1
Bug Hunt Bench22.3Planted Bugs Fixed (out of 105)56
AI Coding Daily (Claude Code) - Total53.69Total points (max 60)50
Bug Hunt Bench - LMS11Planted Bugs Fixed (out of 60)45.6
EnigmaForge39.8Task success (%) over 600 procedurally generated story insta42.9

Interactive version: theaggregate.ai/model?slug=claude-opus-5-5-low · How It Works · Data refreshed daily, snapshot 2026-09-29.