Seizure-Semiology-Suite - Report Generation: leaderboard
Metric: Seizure Report Quality Index, Seizure RQI (0-100): weighted structural completeness (15%), symptom coverage (35%), key localizing features (25%) and temporal-relation F1 (25%) of the generated semiology report against the clinician report, extracted by a Qwen3-Plus LLM, with multiplicative penalties for hallucinated features, off-topic content and excess length and a cap of 50 for hazardous statements; mean over the 82 held-out test videos, 2 FPS sliding windows merged by an LLM, temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Lingshu-32B | 39.8 |
| 2 | Qwen 3 VL 32B Instruct | 39.13 |
| 3 | Qwen 3 VL 8B Instruct | 38.2 |
| 4 | Qwen3 Omni 30B A3B Instruct | 37.52 |
| 5 | Qwen 2.5 VL 32B Instruct | 37.48 |
| 6 | Qwen 2.5 VL 72B Instruct | 37.37 |
| 7 | Qwen 2.5 VL 7B Instruct | 36.94 |
| 8 | Qwen2.5-Omni-7B | 35.91 |
Interactive version: theaggregate.ai/benchmark?slug=seizure-semiology-suite-report-generation · How It Works · Data refreshed daily, snapshot 2026-10-07.