AEGIS (Academic Image Forensics) - Textual Artifact Recognition: leaderboard

Metric: Macro-F1 (%) over two classes, telling whether the text regions of an image carry AI-synthesis traces, on AEGIS academic images (seven categories, 39 subtypes; real panels and forgeries from 25 generative models under four forgery strategies), zero-shot minimal prompts, lossless PNG input; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 26 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)84.67
2GPT-4.182.74
3Gemini 2.5 Flash80.97
4O4 Mini (High)80.02
5Doubao-Seed-1.6 (Thinking)77.57
6Seed-1.676.9
7GPT-5.176.38
8Qwen 2.5 VL 72B Instruct70.76
9Llama 4 Maverick68.32
10Ministral 3 14B62.25
11Doubao-Seed-1.6-Flash60.39
12Claude Sonnet 4.558.7
13Gemma 3 27B (IT)58.39

Interactive version: theaggregate.ai/benchmark?slug=aegis-academic-image-forensics-textual-artifact-recognition · How It Works · Data refreshed daily, snapshot 2026-10-07.