VIABLE - Flawed-Answer Preference: leaderboard
Metric: Gold-pick rate (%; 33,000 gold-versus-failure-injected pairs, judge picks the gold response in both presentation orders, over the WAD, VisAssist and VIA-EgoDex corpora). Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 65.4 |
| 2 | Claude Sonnet 4.6 | 50.2 |
| 3 | Qwen 3 VL 8B | 27.5 |
Interactive version: theaggregate.ai/benchmark?slug=viable-flawed-answer-preference · How It Works · Data refreshed daily, snapshot 2026-09-25.