MedStreamBench - Future: leaderboard
Metric: Content correctness (%; questions about events after the query time, single-turn and streaming, answering unanswerable until evidence suffices; normalized exact match with judge fallback for closed questions, rubric of correctness, grounding and safety for open questions; frames every 2 s within the evidence window; 0-1 x100). Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 43.84 |
| 2 | InternVL3.5-8B | 38.47 |
| 3 | Qwen 2.5 VL 7B | 36.56 |
| 4 | Lingshu-7B | 34.17 |
Interactive version: theaggregate.ai/benchmark?slug=medstreambench-future · How It Works · Data refreshed daily, snapshot 2026-09-29.