Claude Opus 4.8 (Non-reasoning): benchmark results
Provider: Anthropic. Released 2026-05-28. Access: API.
Unified ELO 1753 ± 25, rank #253 of 2075 rated models, from 25 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| StudentBench - Practice Design | 0.96 | Expert tutor pairwise preference, Bradley-Terry score (pract | 83.3 |
| WeirdML | 70.45 | Average Score | 81.4 |
| SteerBench-Work - Pass^5 | 84.9 | Scenarios Correct in All 5 Trials (%) | 79.3 |
| K-Bench (Therapeutic v0) | 97.81 | Overall (%, Therapeutic v0 prompt) | 67.7 |
| ComboShoppingBench - Coupon Legality | 95.9 | Pass rate (%) | 66.7 |
| StudentBench - Lesson Planning | 0.59 | Expert tutor pairwise preference, Bradley-Terry score (lesso | 66.7 |
| K-Bench | 98.11 | Overall (%, Default v1 prompt) | 65.6 |
| ComboShoppingBench - Coupon-ID Validity | 98.6 | Pass rate (%) | 64.3 |
| K-Bench - Risk (Therapeutic v0) | 91.83 | Risk Score (%, Therapeutic v0 prompt) | 62.1 |
| K-Bench - Risk | 92.39 | Risk Score (%, Default v1 prompt) | 59 |
| SteerBench-Work | 86 | Mean Trial Accuracy (%) | 58.6 |
| ComboShoppingBench - Claim Faithfulness | 78.7 | Pass rate (%; LLM-judged) | 54.8 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-04.