AEGIS (Academic Image Forensics) - Manipulation Classification: leaderboard
Metric: Macro-F1 (%) over three classes, naming a highlighted edited region, given the original caption, as an insertion, removal or alteration, on AEGIS academic images (seven categories, 39 subtypes; real panels and forgeries from 25 generative models under four forgery strategies), zero-shot minimal prompts, lossless PNG input; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 26 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.1 | 48.34 |
| 2 | Gemini 3 Pro (Preview) | 45.42 |
| 3 | Doubao-Seed-1.6 (Thinking) | 41.85 |
| 4 | Llama 4 Maverick | 40.18 |
| 5 | GPT-4.1 | 40.05 |
| 6 | Claude Sonnet 4.5 | 39.7 |
| 7 | Gemini 2.5 Flash | 37.28 |
| 8 | Seed-1.6 | 37.08 |
| 9 | Qwen 2.5 VL 72B Instruct | 35.96 |
| 10 | O4 Mini (High) | 35.54 |
| 11 | Gemma 3 27B (IT) | 34.76 |
| 12 | Ministral 3 14B | 33.89 |
| 13 | Doubao-Seed-1.6-Flash | 30.96 |
Interactive version: theaggregate.ai/benchmark?slug=aegis-academic-image-forensics-manipulation-classification · How It Works · Data refreshed daily, snapshot 2026-10-07.