M-Drama - Summary: leaderboard

Metric: Summary judge score (out of 10): DeepSeek-V3.2 rating on a 1-10 rubric of accuracy, completeness and fluency against the human-verified reference description, penalizing hallucinated and wrong facts most; 1,036 episode summaries; test split of 1,036 Chinese and English micro-drama episodes; zero-shot, 192 frames sampled at 2 fps, temperature 0.7, top-p 0.8. Source: arxiv.org. Saturation forecast: Around December 2026. 11 models tracked.

Top models

#ModelScore
1GPT-57.54
2Gemini 3 Flash6.21
3Gemini 2.5 Flash5.78
4Qwen 2.5 VL 72B Instruct5.03
5Qwen 3 VL 32B (Thinking)4.93
6Qwen 3 VL 8B (Thinking)4.18
7Qwen 2.5 VL 7B Instruct3.67

Interactive version: theaggregate.ai/benchmark?slug=m-drama-summary · How It Works · Data refreshed daily, snapshot 2026-09-26.