MemeBench (English): leaderboard

Metric: Success rate (%) on the 485 English memes of 1,253 Chinese and English memes centred on anime, comics, games and online subcultures, explained closed-book from the image; scored against VIKR checklists under the intersection of Gemini-3.1-Pro and GPT-5.1 judges, unscored responses counted as failures. Source: arxiv.org. Saturation forecast: Around December 2026. 26 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)76.7
2Gemini 3 Flash72.6
3Claude Opus 554.2
4GPT-5.6 Sol53.4
5Kimi K2.546.6
6GPT-5.6 Terra36.5
7Qwen 3.6 Plus33
8Claude Sonnet 526.6
9GPT-5.6 Luna22.1
10Qwen 3 VL 235B A22B Instruct21.4
11MiMo-V2.519.8
12MiMo-V2-Omni19.8
13Qwen 3 VL 235B A22B (Thinking)15.7
14Qwen 3.6 Flash15.1
15Qwen 3 VL 32B Instruct12.4

Interactive version: theaggregate.ai/benchmark?slug=memebench-english · How It Works · Data refreshed daily, snapshot 2026-09-29.