AEGIS (Academic Image Forensics): leaderboard

Metric: Normalized Forensic Index (%): harmonic mean of the macro-F1 scores of the three classification tasks and the correct localization accuracy of tampering pinpointing, times the square root of one minus the over-localization rate, on AEGIS academic images (seven categories, 39 subtypes; real panels and forgeries from 25 generative models under four forgery strategies), zero-shot minimal prompts, lossless PNG input; higher is better. Source: arxiv.org. Saturation forecast: Around April 2028. 26 models tracked.

Top models

#ModelScore
1GPT-5.148.8
2Gemini 2.5 Flash47.02
3Doubao-Seed-1.6 (Thinking)46.66
4Gemini 3 Pro (Preview)45.79
5GPT-4.143.31
6O4 Mini (High)42.77
7Seed-1.641.73
8Qwen 2.5 VL 72B Instruct38.71
9Gemma 3 27B (IT)37.55
10Doubao-Seed-1.6-Flash36.03
11Llama 4 Maverick29.15
12Claude Sonnet 4.526.83
13Ministral 3 14B19.9

Interactive version: theaggregate.ai/benchmark?slug=aegis-academic-image-forensics · How It Works · Data refreshed daily, snapshot 2026-10-07.