AHA-Memes: leaderboard

Metric: Binary hateful versus not-hateful macro-F1 (%), zero-shot on the 5K human-annotated gold test split of Arabic memes (image with its OCR text); higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 9 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro71.1
2Qwen 3 VL 8B Instruct64.3
3GPT-562.8
4Gemini 3.5 Flash49.9
5Qwen 3 VL 8B (Thinking)46.9
6InternVL3.5-8B45.8

Interactive version: theaggregate.ai/benchmark?slug=aha-memes · How It Works · Data refreshed daily, snapshot 2026-09-29.