RebusBench (3-Shot) - GPT-4o Judge: leaderboard

Metric: GPT-4o semantic judge score (0-1, times 100) rating the similarity of the predicted answer to the gold answer, averaged over the 1,164 rebus puzzles of RebusBench, three in-context examples in the text prompt; higher is better. Source: arxiv.org. Saturation forecast: Around November 2027. 9 models tracked.

Top models

#ModelScore
1Qwen 2.5 VL 32B Instruct15.08
2Qwen 2.5 VL 7B Instruct14.96

Interactive version: theaggregate.ai/benchmark?slug=rebusbench-3-shot-gpt-4o-judge · How It Works · Data refreshed daily, snapshot 2026-10-07.