AEGIS (Academic Image Forensics): leaderboard
Metric: Normalized Forensic Index (%): harmonic mean of the macro-F1 scores of the three classification tasks and the correct localization accuracy of tampering pinpointing, times the square root of one minus the over-localization rate, on AEGIS academic images (seven categories, 39 subtypes; real panels and forgeries from 25 generative models under four forgery strategies), zero-shot minimal prompts, lossless PNG input; higher is better. Source: arxiv.org. Saturation forecast: Around April 2028. 26 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.1 | 48.8 |
| 2 | Gemini 2.5 Flash | 47.02 |
| 3 | Doubao-Seed-1.6 (Thinking) | 46.66 |
| 4 | Gemini 3 Pro (Preview) | 45.79 |
| 5 | GPT-4.1 | 43.31 |
| 6 | O4 Mini (High) | 42.77 |
| 7 | Seed-1.6 | 41.73 |
| 8 | Qwen 2.5 VL 72B Instruct | 38.71 |
| 9 | Gemma 3 27B (IT) | 37.55 |
| 10 | Doubao-Seed-1.6-Flash | 36.03 |
| 11 | Llama 4 Maverick | 29.15 |
| 12 | Claude Sonnet 4.5 | 26.83 |
| 13 | Ministral 3 14B | 19.9 |
Interactive version: theaggregate.ai/benchmark?slug=aegis-academic-image-forensics · How It Works · Data refreshed daily, snapshot 2026-10-07.