M2-Verify-Gen: leaderboard

Metric: Macro-F1 (%) of binary claim consistency verification (supported or refuted) on the M2-Verify-Gen test set (figures and captions from 16 arXiv categories, negation-perturbed claims), zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 8 models tracked.

Top models

#ModelScore
1Qwen 2.5 VL 7B Instruct68.1
2InternVL3-8B63.8
3Mistral Small 3.163.1
4GPT-5 Mini59.9
5GPT-4o Mini56.6
6Phi-4 Multimodal Instruct45.7
7Pixtral-12B41.9

Interactive version: theaggregate.ai/benchmark?slug=m2-verify-gen · How It Works · Data refreshed daily, snapshot 2026-10-07.