M-Drama - Multiple Choice: leaderboard
Metric: Pass@1 accuracy (%; exact match on the option letter over the 1,562 multiple-choice questions in nine character, plot and scene task types, three sampled responses per question; test split of 1,036 Chinese and English micro-drama episodes; zero-shot, 192 frames sampled at 2 fps, temperature 0.7, top-p 0.8). Source: arxiv.org. Saturation forecast: Around December 2026. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash | 85.3 |
| 2 | GPT-5 | 83.7 |
| 3 | Gemini 2.5 Flash | 76.7 |
| 4 | Qwen 3 VL 32B (Thinking) | 66.5 |
| 5 | Qwen 3 VL 8B (Thinking) | 56.5 |
| 6 | Qwen 2.5 VL 72B Instruct | 54.6 |
| 7 | Qwen 2.5 VL 7B Instruct | 38.6 |
Interactive version: theaggregate.ai/benchmark?slug=m-drama-multiple-choice · How It Works · Data refreshed daily, snapshot 2026-09-26.