M2-Verify-Med: leaderboard

Metric: Macro-F1 (%) of binary claim consistency verification (supported or refuted) on the M2-Verify-Med test set (PubMed figures and captions, claims with expert-curated medical perturbations), zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 10 models tracked.

Top models

#ModelScore
1GPT-4o Mini73
2Mistral Small 3.171.4
3Qwen 2.5 VL 7B Instruct70.4
4InternVL3-8B69.6
5Pixtral-12B69.6
6GPT-5 Mini67.2
7Phi-4 Multimodal Instruct65.8

Interactive version: theaggregate.ai/benchmark?slug=m2-verify-med · How It Works · Data refreshed daily, snapshot 2026-10-07.