M-Drama - Summary: leaderboard
Metric: Summary judge score (out of 10): DeepSeek-V3.2 rating on a 1-10 rubric of accuracy, completeness and fluency against the human-verified reference description, penalizing hallucinated and wrong facts most; 1,036 episode summaries; test split of 1,036 Chinese and English micro-drama episodes; zero-shot, 192 frames sampled at 2 fps, temperature 0.7, top-p 0.8. Source: arxiv.org. Saturation forecast: Around December 2026. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 7.54 |
| 2 | Gemini 3 Flash | 6.21 |
| 3 | Gemini 2.5 Flash | 5.78 |
| 4 | Qwen 2.5 VL 72B Instruct | 5.03 |
| 5 | Qwen 3 VL 32B (Thinking) | 4.93 |
| 6 | Qwen 3 VL 8B (Thinking) | 4.18 |
| 7 | Qwen 2.5 VL 7B Instruct | 3.67 |
Interactive version: theaggregate.ai/benchmark?slug=m-drama-summary · How It Works · Data refreshed daily, snapshot 2026-09-26.