STRAND: leaderboard

Metric: Faithful accuracy (%; share of the 977 target yes/no questions over 88 videos for which the model answers the target and every prerequisite sub-question correctly, computed over all targets; 2,516 sub-questions; end-to-end video input; majority-class predictor 19.1). Source: arxiv.org. Saturation forecast: Around 2034. 11 models tracked.

Top models

#ModelScore
1Qwen 3 VL 32B (Thinking)46.1
2Qwen 3.5 27B45.5
3Gemini 3.1 Pro (Preview)38.9
4Claude Sonnet 4.630.2
5InternVL3-78B29.7
6Gemini 3 Flash28.7
7GPT-520

Interactive version: theaggregate.ai/benchmark?slug=strand · How It Works · Data refreshed daily, snapshot 2026-09-26.