AHA-Memes - Fine-Grained Hate Types: leaderboard
Metric: Fine-grained macro-F1 (%) over ten hate and non-hate type labels (multi-label), zero-shot on the 5K human-annotated gold test split of Arabic memes (image with its OCR text); higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 34 |
| 2 | GPT-5 | 30.1 |
| 3 | Gemini 3.5 Flash | 27.1 |
| 4 | Qwen 3 VL 8B (Thinking) | 19.1 |
| 5 | Qwen 3 VL 8B Instruct | 17.6 |
| 6 | InternVL3.5-8B | 16.4 |
Interactive version: theaggregate.ai/benchmark?slug=aha-memes-fine-grained-hate-types · How It Works · Data refreshed daily, snapshot 2026-09-29.