PaSBench-Video - Healthcare: leaderboard

Metric: First-Act@1.0 (%) on the healthcare (33 risk videos, synthesized clips): share of risk videos whose first warning, issued while the model streams 3-second windows of past frames only, lands between the annotated risk onset and 1 s before the accident and names the correct risk source; temperature 0, identical prompts; false positives on the no-risk split are not part of this score; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 13 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.627.3
2Qwen 3.5 27B18.2
3GPT-5.412.1
4Claude Opus 4.712.1
5Qwen 3.5 35B A3B12.1
6Qwen 3.5 4B12.1
7Qwen 3.5 9B9.1
8GPT-5.4 Mini9.1
9Gemma 4 26B6.1
10Gemma 4 31B3
11Claude Haiku 4.50

Interactive version: theaggregate.ai/benchmark?slug=pasbench-video-healthcare · How It Works · Data refreshed daily, snapshot 2026-09-29.