MEDit-Bench - F1@0.5: leaderboard
Metric: F1@0.5 (%; cut-level F1 after one-to-one matching of predicted and reference cuts at temporal IoU of at least 0.5, averaged over the 540 video, message and reference-edit triplets (60 long-form videos, three editing messages each, three professional edits per message); zero-shot: the model maps the video and the message to a sequence of cuts). Source: arxiv.org. Saturation forecast: Around 2030. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro | 15.2 |
| 2 | GPT-5 | 15.1 |
| 3 | Qwen 3 VL 32B | 12.6 |
| 4 | GPT-5 Mini | 11.2 |
| 5 | Gemini 2.5 Pro | 10.2 |
| 6 | Gemma 4 E4B (Non-reasoning) | 7.1 |
| 7 | InternVL3.5-8B | 4.2 |
Interactive version: theaggregate.ai/benchmark?slug=medit-bench-f1-0-5 · How It Works · Data refreshed daily, snapshot 2026-09-29.