AutoMedBench: leaderboard

48 medical-imaging tasks in 7 tracks (segmentation, VQA, report generation and more) that a coding agent must plan, run and submit; UC Santa Cruz and NVIDIA (2026); half rubric, half task metric.

Metric: Average Overall Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 7 models tracked.

Top models

#ModelScore
1Claude Opus 4.669.69
2GLM-563.32
3Gemini 3.1 Pro (Preview)61.23
4MiniMax-M2.552.34
5Kimi K2.533.86

Interactive version: theaggregate.ai/benchmark?slug=automedbench · How It Works · Data refreshed daily, snapshot 2026-09-05.