MedStreamBench - Present: leaderboard
Metric: Content correctness (%; questions about the visual state near the query time; normalized exact match with judge fallback for closed questions, rubric of correctness, grounding and safety for open questions; frames every 2 s within the evidence window; 0-1 x100). Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 41.4 |
| 2 | InternVL3.5-8B | 34.61 |
| 3 | Qwen 2.5 VL 7B | 34.46 |
| 4 | Lingshu-7B | 33.02 |
Interactive version: theaggregate.ai/benchmark?slug=medstreambench-present · How It Works · Data refreshed daily, snapshot 2026-09-29.