Seizure-Semiology-Suite - Report Generation: leaderboard

Metric: Seizure Report Quality Index, Seizure RQI (0-100): weighted structural completeness (15%), symptom coverage (35%), key localizing features (25%) and temporal-relation F1 (25%) of the generated semiology report against the clinician report, extracted by a Qwen3-Plus LLM, with multiplicative penalties for hallucinated features, off-topic content and excess length and a cap of 50 for hazardous statements; mean over the 82 held-out test videos, 2 FPS sliding windows merged by an LLM, temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 10 models tracked.

Top models

#ModelScore
1Lingshu-32B39.8
2Qwen 3 VL 32B Instruct39.13
3Qwen 3 VL 8B Instruct38.2
4Qwen3 Omni 30B A3B Instruct37.52
5Qwen 2.5 VL 32B Instruct37.48
6Qwen 2.5 VL 72B Instruct37.37
7Qwen 2.5 VL 7B Instruct36.94
8Qwen2.5-Omni-7B35.91

Interactive version: theaggregate.ai/benchmark?slug=seizure-semiology-suite-report-generation · How It Works · Data refreshed daily, snapshot 2026-10-07.