PaSBench-Video - Healthcare: leaderboard
Metric: First-Act@1.0 (%) on the healthcare (33 risk videos, synthesized clips): share of risk videos whose first warning, issued while the model streams 3-second windows of past frames only, lands between the annotated risk onset and 1 s before the accident and names the correct risk source; temperature 0, identical prompts; false positives on the no-risk split are not part of this score; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.6 | 27.3 |
| 2 | Qwen 3.5 27B | 18.2 |
| 3 | GPT-5.4 | 12.1 |
| 4 | Claude Opus 4.7 | 12.1 |
| 5 | Qwen 3.5 35B A3B | 12.1 |
| 6 | Qwen 3.5 4B | 12.1 |
| 7 | Qwen 3.5 9B | 9.1 |
| 8 | GPT-5.4 Mini | 9.1 |
| 9 | Gemma 4 26B | 6.1 |
| 10 | Gemma 4 31B | 3 |
| 11 | Claude Haiku 4.5 | 0 |
Interactive version: theaggregate.ai/benchmark?slug=pasbench-video-healthcare · How It Works · Data refreshed daily, snapshot 2026-09-29.