MedStreamBench - Retrospective: leaderboard

Metric: Content correctness (%; questions about events within the evidence window ending at the query time; normalized exact match with judge fallback for closed questions, rubric of correctness, grounding and safety for open questions; frames every 2 s within the evidence window; 0-1 x100). Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro34.28
2InternVL3.5-8B25.36
3Qwen 2.5 VL 7B24.13
4Lingshu-7B23.29

Interactive version: theaggregate.ai/benchmark?slug=medstreambench-retrospective · How It Works · Data refreshed daily, snapshot 2026-09-29.