HELM Reasoning - Reasoning Scenarios: leaderboard
Metric: Mean score. Source: crfm.stanford.edu. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 Mini (2025-08-07) | 83.51 |
| 2 | Qwen 3 4B | 70.62 |
| 3 | Qwen 3 8B | 67.69 |
| 4 | Qwen 3 1.7B | 48.24 |
| 5 | Qwen 3 0.6B | 24.62 |
Interactive version: theaggregate.ai/benchmark?slug=helm-reasoning-reasoning-scenarios · How It Works · Data refreshed daily, snapshot 2026-09-08.