SynCred-Bench: leaderboard

Metric: True positive rate (%) on the 600 AI-generated credible-form images at an operating point whose false positive rate on the 450 real FP450 images is at most 5% (threshold swept over the stated confidence when the default FPR exceeds 5%); MLLM-as-judge detection prompt, image only. Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1Claude Opus 4.669.5
2Claude Sonnet 4.655.3
3Gemini 3.1 Pro (Preview)12.7
4GPT-5.411.5
5Qwen 3.6 Plus3.7
6Qwen 3 VL 32B Instruct2.2
7Gemini 3.1 Flash Lite1.5
8GLM-5V Turbo0.5
9Grok 4.30.3
10GPT-4o0
11Qwen 3.5 9B0
12Qwen 3.5 35B A3B0
13Llama 3.2 11B Instruct0
14Qwen 3 VL 8B Instruct0
15Pixtral Large0

Interactive version: theaggregate.ai/benchmark?slug=syncred-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.