CT-Bench (Lesion VQA): leaderboard
Metric: Accuracy (%) pooled over the questions of all 13 task settings (seven tasks with the lesion bounding box drawn, six without; the paper's Average Total, a question-weighted mean of the task cells) on the CT-Bench QA test split (four-option multiple choice on DeepLesion CT: lesion description from one slice (200 questions per box setting) or nine consecutive slices (200), text-to-slice retrieval (200), box localization (100, with box only), size (100) and attribute classification from one or nine slices (100 each)); distractors are hard negatives retrieved with BiomedCLIP and checked by a physician; zero-shot; chance is 25%; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini (CT-Bench checkpoint unspecified) | 33.8 | |
| 2 | Llama-3-8B-Dragonfly-Med-v1 | 31.7 | |
| 3 | GPT-4V | 28.4 | |
| 4 | RadFM | 15.3 | |
| 5 | LLaVA-Med | 14.5 |
Interactive version: theaggregate.ai/benchmark?slug=ct-bench-lesion-vqa · How It Works · Data refreshed daily, snapshot 2026-10-11.