AEGIS (Academic Image Forensics) - Forgery Scope Discrimination: leaderboard
Metric: Macro-F1 (%) over three classes, telling whether an image is authentic, partly AI-edited or entirely AI-generated (answers of Not Sure, an offered abstention, are left out), on AEGIS academic images (seven categories, 39 subtypes; real panels and forgeries from 25 generative models under four forgery strategies), zero-shot minimal prompts, lossless PNG input; higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 26 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4.1 | 59.98 |
| 2 | Gemini 2.5 Flash | 53.78 |
| 3 | Gemini 3 Pro (Preview) | 53.44 |
| 4 | O4 Mini (High) | 49.35 |
| 5 | GPT-5.1 | 46.16 |
| 6 | Gemma 3 27B (IT) | 42.36 |
| 7 | Doubao-Seed-1.6 (Thinking) | 37.13 |
| 8 | Llama 4 Maverick | 37 |
| 9 | Seed-1.6 | 35.38 |
| 10 | Qwen 2.5 VL 72B Instruct | 33.16 |
| 11 | Ministral 3 14B | 25.44 |
| 12 | Doubao-Seed-1.6-Flash | 23.77 |
| 13 | Claude Sonnet 4.5 | 19.62 |
Interactive version: theaggregate.ai/benchmark?slug=aegis-academic-image-forensics-forgery-scope-discrimination · How It Works · Data refreshed daily, snapshot 2026-10-07.