MemeBridge: leaderboard

Metric: Performance score (out of 5) summing the cosine similarity of the generated explanation to the original and to the GPT-4-rewritten explanation plus the multiple-choice, sentiment and emotion accuracies scaled so that each chance-level task scores about 1 (621 U.S.-originated memes with U.S. crowd labels, default prompt (no role-play), the same test given to Chinese participants). Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.

Top models

#ModelScore
1GPT-4o3.24

Interactive version: theaggregate.ai/benchmark?slug=memebridge · How It Works · Data refreshed daily, snapshot 2026-09-26.