CT-Bench (Lesion VQA) - With Box: leaderboard
Metric: Accuracy (%) pooled over the questions of the seven tasks posed with the lesion bounding box drawn on the image (question-weighted mean of the task cells), on the CT-Bench QA test split (four-option multiple choice on DeepLesion CT: lesion description from one slice (200 questions per box setting) or nine consecutive slices (200), text-to-slice retrieval (200), box localization (100, with box only), size (100) and attribute classification from one or nine slices (100 each)); distractors are hard negatives retrieved with BiomedCLIP and checked by a physician; zero-shot; chance is 25%; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini (CT-Bench checkpoint unspecified) | 36.1 | |
| 2 | Llama-3-8B-Dragonfly-Med-v1 | 31.1 | |
| 3 | GPT-4V | 28.2 | |
| 4 | RadFM | 14.9 | |
| 5 | LLaVA-Med | 13.6 |
Interactive version: theaggregate.ai/benchmark?slug=ct-bench-lesion-vqa-with-box · How It Works · Data refreshed daily, snapshot 2026-10-11.