Claude 2 — benchmark results

Anthropic Claude 2 legacy chat model. Provider: Anthropic. Released 2023-07-11. Access: API.

Unified ELO 1527 ± 25, rank #694 of 1776 rated models, from 59 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Trustworthy - Toxicity92.11Trust Score (%)100
LLM Trustworthy Leaderboard84.52Average Trust Score (%)100
NPHardEval - TSP D87.27Accuracy (%)100
SALAD-Bench93.91Average Safety Score (%)100
SALAD-Bench Attack88.06Safety Score (%)100
SALAD-Bench Base99.77Safety Score (%)100
TriviaQA87.5Accuracy (%)97.4
LLM Trustworthy - Stereotype100Trust Score (%)94
LLM Trustworthy - Out-of-Distribution85.77Trust Score (%)92
AgentBoard48.9Progress Rate (self-reported)91.7
NPHardEval26.42Average Accuracy (%)90.9
NPHardEval - SPP37.27Accuracy (%)90.9

Interactive version: theaggregate.ai/model?slug=claude-2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.