Claude Opus 5.5 (Low): benchmark results
Provider: Anthropic. Released 2026-09-22. Access: API.
Unified ELO 1884 ± 28, rank #46 of 2055 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EnigmaForge - Fact F1 | 91.9 | World-fact recovery F1 (%) over 600 instances, all items | 95.2 |
| DuelLab Overall | 59.4 | DuelLab Score | 90 |
| DataBench | 54 | Score (%) | 85.7 |
| ARC-AGI-2 | 70.14 | Accuracy (%) | 80.3 |
| ARC-AGI-1 | 88.5 | Accuracy (%) | 72.5 |
| Bug Hunt Bench - VS Code Extension | 11.3 | Planted Bugs Fixed (out of 45) | 63.7 |
| EnigmaForge - Intuition | 22 | Task success (%) on implicit-condition instances, where the | 61.9 |
| NonoBench | 60 | Overall Accuracy (%) | 60.1 |
| Bug Hunt Bench | 22.3 | Planted Bugs Fixed (out of 105) | 56 |
| AI Coding Daily (Claude Code) - Total | 53.69 | Total points (max 60) | 50 |
| Bug Hunt Bench - LMS | 11 | Planted Bugs Fixed (out of 60) | 45.6 |
| EnigmaForge | 39.8 | Task success (%) over 600 procedurally generated story insta | 42.9 |
Interactive version: theaggregate.ai/model?slug=claude-opus-5-5-low · How It Works · Data refreshed daily, snapshot 2026-09-29.