CT-Bench (Lesion VQA) - With Box: leaderboard

Metric: Accuracy (%) pooled over the questions of the seven tasks posed with the lesion bounding box drawn on the image (question-weighted mean of the task cells), on the CT-Bench QA test split (four-option multiple choice on DeepLesion CT: lesion description from one slice (200 questions per box setting) or nine consecutive slices (200), text-to-slice retrieval (200), box localization (100, with box only), size (100) and attribute classification from one or nine slices (100 each)); distractors are hard negatives retrieved with BiomedCLIP and checked by a physician; zero-shot; chance is 25%; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.

Top models

#ModelScoreOverall rank
1Gemini (CT-Bench checkpoint unspecified)36.1
2Llama-3-8B-Dragonfly-Med-v131.1
3GPT-4V28.2
4RadFM14.9
5LLaVA-Med13.6

Interactive version: theaggregate.ai/benchmark?slug=ct-bench-lesion-vqa-with-box · How It Works · Data refreshed daily, snapshot 2026-10-11.