MEDit-Bench - F1@0.5: leaderboard

Metric: F1@0.5 (%; cut-level F1 after one-to-one matching of predicted and reference cuts at temporal IoU of at least 0.5, averaged over the 540 video, message and reference-edit triplets (60 long-form videos, three editing messages each, three professional edits per message); zero-shot: the model maps the video and the message to a sequence of cuts). Source: arxiv.org. Saturation forecast: Around 2030. 9 models tracked.

Top models

#ModelScore
1Gemini 3 Pro15.2
2GPT-515.1
3Qwen 3 VL 32B12.6
4GPT-5 Mini11.2
5Gemini 2.5 Pro10.2
6Gemma 4 E4B (Non-reasoning)7.1
7InternVL3.5-8B4.2

Interactive version: theaggregate.ai/benchmark?slug=medit-bench-f1-0-5 · How It Works · Data refreshed daily, snapshot 2026-09-29.