Claude 3.7: benchmark results

Provider: Anthropic. Released 2025-02-24. Access: API.

Unified ELO 1678 ± 31, rank #377 of 2656 rated models, from 18 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EduGuardBench77RFS (self-reported)100
EComStage - Query Rewrite81.99Task Score (%)96.9
EComStage - RAG-QA69.38Task Score (%)78.1
GosuEvals70.6Score (%)77.8
EComStage - Intent Recognition88Task Score (%)75
EComStage - Query Match98.39Task Score (%)70.3
KORGym - Puzzle0.55Score61.1
EComStage - Scenario Route82.32Task Score (%)59.4
EComStage81.89Task Score (%)53.1
KORGym0.35Score44.4
KORGym - Mathematical and Logical0.26Score44.4
KORGym - Control and Interaction0.25Score41.7

Interactive version: theaggregate.ai/model?slug=claude-3-7 · How It Works · Data refreshed daily, snapshot 2026-09-19.