EviPathBench - Image-to-Text Matching: leaderboard

Metric: Accuracy (%) pooled over all questions (the All part of each cell) on Task 1: choosing the pathologist's finding for a slide region image among four descriptions; EviPathBench: pathologist-authored diagnostic paths on 1,822 TCGA whole-slide images (16 organs), findings at 2.5x, 10x and 40x magnification; four options per question with distractors mined from the same organ and magnification; parsed choice scored by exact match; chance 25; higher is better. Source: arxiv.org. Saturation forecast: Around February 2027. 19 models tracked.

Top models

#ModelScore
1Gemini 3 Flash63.48
2Kimi K2.558.08
3Qwen 3.5 Flash54.36
4Lingshu-7B53.23
5GPT-5.252.23
6Lingshu-32B47.36
7MedGemma 1.5 4B36.37
8Qwen 2.5 VL 7B Instruct33.22
9MedGemma-4B31.63
10MedGemma-27B-IT27.91

Interactive version: theaggregate.ai/benchmark?slug=evipathbench-image-to-text-matching · How It Works · Data refreshed daily, snapshot 2026-09-29.