VIABLE - Flawed-Answer Preference: leaderboard

Metric: Gold-pick rate (%; 33,000 gold-versus-failure-injected pairs, judge picks the gold response in both presentation orders, over the WAD, VisAssist and VIA-EgoDex corpora). Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.

Top models

#ModelScore
1GPT-5.465.4
2Claude Sonnet 4.650.2
3Qwen 3 VL 8B27.5

Interactive version: theaggregate.ai/benchmark?slug=viable-flawed-answer-preference · How It Works · Data refreshed daily, snapshot 2026-09-25.