HELM VHELM - Exams V - Arabic (Natural Science): leaderboard

Metric: Prefix Quasi-Exact Match (%). Source: crfm.stanford.edu. 48 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro (Preview 03-25)92.86
2O3 (2025-04-16)90.48
3O4 Mini (2025-04-16)89.29
4O1 (2024-12-17)88.1
5GPT-4.572.62
6Gemini 2.0 Flash (Preview)71.43
7Claude 3.5 Sonnet (20240620)66.67
8Gemini 2.0 Flash66.67
9GPT-4.1 (2025-04-14)66.67
10GPT-4o (2024-08-06)65.48
11GPT-4o (2024-11-20)63.1
12GPT-4o (2024-05-13)61.9
13Gemini 2.0 Pro (Preview 02-05)60.71
14Claude 3.5 Sonnet (20241022)58.33
15Gemini 2.0 Flash Lite58.33

Interactive version: theaggregate.ai/benchmark?slug=helm-vhelm-exams-v-arabic-natural-science · How It Works · Data refreshed daily, snapshot 2026-09-19.