RebusBench (1-Shot) - GPT-4o Judge: leaderboard
Metric: GPT-4o semantic judge score (0-1, times 100) rating the similarity of the predicted answer to the gold answer, averaged over the 1,164 rebus puzzles of RebusBench, one in-context example in the text prompt; higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 2.5 VL 32B Instruct | 14.78 |
| 2 | Qwen 2.5 VL 7B Instruct | 14.13 |
Interactive version: theaggregate.ai/benchmark?slug=rebusbench-1-shot-gpt-4o-judge · How It Works · Data refreshed daily, snapshot 2026-10-07.