AEGIS (Academic Image Forensics) - Forgery Scope Discrimination: leaderboard

Metric: Macro-F1 (%) over three classes, telling whether an image is authentic, partly AI-edited or entirely AI-generated (answers of Not Sure, an offered abstention, are left out), on AEGIS academic images (seven categories, 39 subtypes; real panels and forgeries from 25 generative models under four forgery strategies), zero-shot minimal prompts, lossless PNG input; higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 26 models tracked.

Top models

#ModelScore
1GPT-4.159.98
2Gemini 2.5 Flash53.78
3Gemini 3 Pro (Preview)53.44
4O4 Mini (High)49.35
5GPT-5.146.16
6Gemma 3 27B (IT)42.36
7Doubao-Seed-1.6 (Thinking)37.13
8Llama 4 Maverick37
9Seed-1.635.38
10Qwen 2.5 VL 72B Instruct33.16
11Ministral 3 14B25.44
12Doubao-Seed-1.6-Flash23.77
13Claude Sonnet 4.519.62

Interactive version: theaggregate.ai/benchmark?slug=aegis-academic-image-forensics-forgery-scope-discrimination · How It Works · Data refreshed daily, snapshot 2026-10-07.