RebusBench (1-Shot) - GPT-4o Judge: leaderboard

Metric: GPT-4o semantic judge score (0-1, times 100) rating the similarity of the predicted answer to the gold answer, averaged over the 1,164 rebus puzzles of RebusBench, one in-context example in the text prompt; higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 9 models tracked.

Top models

#ModelScore
1Qwen 2.5 VL 32B Instruct14.78
2Qwen 2.5 VL 7B Instruct14.13

Interactive version: theaggregate.ai/benchmark?slug=rebusbench-1-shot-gpt-4o-judge · How It Works · Data refreshed daily, snapshot 2026-10-07.