Claude v1.3: benchmark results

Provider: Anthropic. Access: API.

Unified ELO 1487 ± 34, rank #1198 of 2656 rated models, from 31 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Lite - WMT 2014 (ru-en)28.03BLEU-4 (%)94.4
HELM WMT 201421.85BLEU-4 (%)80
HELM v2 Lite - MedQA75Exact Match (%)80
HELM Lite - WMT 2014 (cs-en)23.1BLEU-4 (%)78.9
HELM Lite - WMT 2014 (de-en)20.21BLEU-4 (%)73.3
HELM Lite - WMT 2014 (fr-en)22.75BLEU-4 (%)70.8
HELM Lite - WMT 2014 (hi-en)15.17BLEU-4 (%)66.7
HELM NaturalQuestions (Closed)40.89F1 (%)66.7
HELM Lite - LegalBench - Function of Decision Section41.69Exact Match (%)65.6
HELM Lite57.19Mean win rate (self-reported)64.5
HELM Lite - MATH Level 1 - Prealgebra82.56Equivalent (CoT) (%)63.9
HELM Lite - LegalBench - Abercrombie65.26Exact Match (%)61.1

Interactive version: theaggregate.ai/model?slug=claude-v1-3 · How It Works · Data refreshed daily, snapshot 2026-09-19.