HELM Capabilities - WildBench — leaderboard

Metric: WB Score. Source: crfm.stanford.edu. 51 models tracked.

Top models

#ModelScore
1Kimi K286.19
2O3 (2025-04-16)86.07
3Gemini 2.5 Pro (Preview 03-25)85.67
4GPT-4.1 (2025-04-14)85.43
5O4 Mini (2025-04-16)85.41
6Grok 384.94
7GPT-4.1 Mini83.8
8Claude Opus 4 (20250514)83.3
9DeepSeek V383.05
10GPT-4o (2024-11-20)82.84
11DeepSeek R1 052882.81
12Qwen3 235B A22B FP8 Throughput82.78
13Claude Sonnet 4 (20250514)82.47
14Gemini 2.5 Flash (Preview 04-17)81.72
15Claude 3.7 Sonnet (20250219)81.44

Interactive version: theaggregate.ai/benchmark?slug=helm-capabilities-wildbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.