EVID-Bench - Synthetic Insertion: leaderboard

Metric: Point-level accuracy (%) on the 24 synthetic insertion videos (AI generation: a generated clip inserted into real footage), using the retrieval-augmented verification pipeline of the paper (chain-of-thought analysis, up to six YouTube search-verify-reflect rounds, 64 frames at 480p, temperature 0); each video has 3 to 5 annotated misinformation points, matched by a majority of three LLM judges; mean of three runs; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 9 models tracked.

Top models

#ModelScore
1GPT-5.516.67
2Gemini 3.1 Pro (Preview)13.9
3Qwen 3.5 Plus13.9
4Claude Opus 4.613.89
5Claude Sonnet 4.612.5
6Qwen 3 VL 235B A22B Instruct12.5
7Gemini 3 Flash11.1
8GPT-5.49.72
9GPT-5.4 Mini6.94

Interactive version: theaggregate.ai/benchmark?slug=evid-bench-synthetic-insertion · How It Works · Data refreshed daily, snapshot 2026-09-29.