MedSNIP-Bench: leaderboard

Metric: False-class F1 (%; F1 on false snippets among the 2,524 human-annotated snippets of 276 consumer-health and clinical-vignette answers, 13.6% false; each snippet verified on its own from parametric knowledge (claim-only mode); same prompt for every verifier). Source: arxiv.org. Saturation forecast: Around 2032. 6 models tracked.

Top models

#ModelScore
1GPT-5.4 (High)43.1
2GPT-4o40.1
3GPT-OSS-20B40.1
4Gemma 4 31B (IT)38.9
5Llama 3.3 70B Instruct33.1
6Llama 3.1 8B Instruct27.8

Interactive version: theaggregate.ai/benchmark?slug=medsnip-bench · How It Works · Data refreshed daily, snapshot 2026-09-26.