MedRedFlag - False Assumptions Accommodated: leaderboard

Metric: Share of 100 questions (%) whose response still provides the requested information, implicitly treating the false assumption as valid (GPT-5 judge against a physician-written criterion; lower is better). Source: arxiv.org. Saturation forecast: Around 2029. 5 models tracked.

Top models

#ModelScore
1Claude Opus 4.560
2GPT-573
3Llama 3.3 70B Instruct74
4MedGemma-27B-IT74
5Qwen 3 32B80

Interactive version: theaggregate.ai/benchmark?slug=medredflag-false-assumptions-accommodated · How It Works · Data refreshed daily, snapshot 2026-09-25.