MedStreamBench - Retrospective: leaderboard
Metric: Content correctness (%; questions about events within the evidence window ending at the query time; normalized exact match with judge fallback for closed questions, rubric of correctness, grounding and safety for open questions; frames every 2 s within the evidence window; 0-1 x100). Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 34.28 |
| 2 | InternVL3.5-8B | 25.36 |
| 3 | Qwen 2.5 VL 7B | 24.13 |
| 4 | Lingshu-7B | 23.29 |
Interactive version: theaggregate.ai/benchmark?slug=medstreambench-retrospective · How It Works · Data refreshed daily, snapshot 2026-09-29.