EviPathBench - Text-to-Image Retrieval: leaderboard

Metric: Accuracy (%) pooled over all questions (the All part of each cell) on Task 2: choosing the slide region image that matches a pathologist's finding among four images; EviPathBench: pathologist-authored diagnostic paths on 1,822 TCGA whole-slide images (16 organs), findings at 2.5x, 10x and 40x magnification; four options per question with distractors mined from the same organ and magnification; parsed choice scored by exact match; chance 25; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 16 models tracked.

Top models

#ModelScore
1Gemini 3 Flash67.74
2GPT-5.260.96
3Qwen 3.5 Flash60.05
4Lingshu-7B54.49
5Lingshu-32B44.22
6Qwen 2.5 VL 7B Instruct37
7MedGemma-4B33
8MedGemma 1.5 4B30.89
9MedGemma-27B-IT23.76

Interactive version: theaggregate.ai/benchmark?slug=evipathbench-text-to-image-retrieval · How It Works · Data refreshed daily, snapshot 2026-09-29.