HarmVideoBench - Beyond-Clip Reasoning: leaderboard

Metric: Beyond-Clip Reasoning accuracy (%) on four-option multiple-choice questions about 1,379 real harmful videos (one question per video per level), temperature 0 greedy decoding; API models take the video and audio natively, open models 12 sampled key frames plus the transcript; chance 25; questions that need context beyond the clip; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 19 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)65.8
2GPT-5.563.9
3Claude Opus 4.763.2
4Claude Sonnet 4.662.7
5Qwen 3 VL 235B A22B (Thinking)62.5
6Qwen 3 VL 32B (Thinking)62.1
7Qwen 3 VL 235B A22B Instruct62
8Qwen 3 VL 30B A3B (Thinking)61.2
9Qwen 3 VL 32B Instruct61.1
10Qwen 3 VL 8B (Thinking)59.5
11Qwen 3 VL 30B A3B Instruct58.5
12LFM2.5-VL-1.6B51.4
13GLM-4.5V (Non-reasoning)50.9
14Qwen 3 VL 8B Instruct48.1

Interactive version: theaggregate.ai/benchmark?slug=harmvideobench-beyond-clip-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-29.