HAL AssistantBench — leaderboard

Princeton HAL leaderboard for AssistantBench: web assistance tasks requiring multi-step browsing, information gathering, and synthesis.

Metric: Accuracy (%). Source: hal.cs.princeton.edu. Status: saturation imminent. 15 models tracked.

Top models

#ModelScore
1O3 (Medium)38.81
2GPT-5 (Medium)35.23
3O4 Mini (Low)28.05
4O4 Mini (High)23.84
5GPT-4.1 (2025-04-14)17.39
6Claude 3.7 Sonnet (20250219)16.69
7Claude Opus 4.1 (High)13.75
8Claude 3.7 Sonnet (High)13.08
9Claude Sonnet 4.5 (High)11.8
10DeepSeek R1 05288.75
11Claude Opus 4.1 (20250805)7.26
12Claude Sonnet 4.57.09
13Gemini 2.0 Flash (001)2.62
14DeepSeek V3 (0324)2.03
15DeepSeek R10

Interactive version: theaggregate.ai/benchmark?slug=hal-assistantbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.