HELM AIR-Bench 2024 - #26.12: Misrepresenting voting qualifications: leaderboard

Metric: Refusal Rate (%). Source: crfm.stanford.edu. 87 models tracked.

Top models

#ModelScore
1Llama 3.1 8B Instruct100
2GPT-4.1 Mini100
3GPT-5100
4Claude 3.5 Sonnet (20241022)100
5Qwen 2.5 72B Instruct100
6Claude Sonnet 4.5100
7Claude 3.5 Sonnet (20240620)100
8GPT-5 Nano100
9GPT-OSS-120B100
10Gemini 2.0 Flash100
11Gemini 2.5 Pro (Preview 03-25)100
12Gemini 3 Pro (Preview)100
13GPT-OSS-20B100
14Llama 3 70B Instruct100
15Qwen 3 Next 80B A3B (Thinking)100

Interactive version: theaggregate.ai/benchmark?slug=helm-air-bench-2024-26-12-misrepresenting-voting-qualifications · How It Works · Data refreshed daily, snapshot 2026-09-19.