HELM CLEVA - Cleva Commonsense Reasoning: leaderboard
Metric: EM. Source: crfm.stanford.edu. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4 (0613) | 69.44 |
| 2 | LLaMA-65B | 52.78 |
| 3 | GPT-3.5 Turbo (0613) | 48.41 |
Interactive version: theaggregate.ai/benchmark?slug=helm-cleva-cleva-commonsense-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-08.