Claude 2 — benchmark results
Anthropic Claude 2 legacy chat model. Provider: Anthropic. Released 2023-07-11. Access: API.
Unified ELO 1527 ± 25, rank #694 of 1776 rated models, from 59 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Trustworthy - Toxicity | 92.11 | Trust Score (%) | 100 |
| LLM Trustworthy Leaderboard | 84.52 | Average Trust Score (%) | 100 |
| NPHardEval - TSP D | 87.27 | Accuracy (%) | 100 |
| SALAD-Bench | 93.91 | Average Safety Score (%) | 100 |
| SALAD-Bench Attack | 88.06 | Safety Score (%) | 100 |
| SALAD-Bench Base | 99.77 | Safety Score (%) | 100 |
| TriviaQA | 87.5 | Accuracy (%) | 97.4 |
| LLM Trustworthy - Stereotype | 100 | Trust Score (%) | 94 |
| LLM Trustworthy - Out-of-Distribution | 85.77 | Trust Score (%) | 92 |
| AgentBoard | 48.9 | Progress Rate (self-reported) | 91.7 |
| NPHardEval | 26.42 | Average Accuracy (%) | 90.9 |
| NPHardEval - SPP | 37.27 | Accuracy (%) | 90.9 |
Interactive version: theaggregate.ai/model?slug=claude-2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.