A2RBench - Symbolic Tasks (Remapped Symbols): leaderboard
Metric: Accuracy (%) on the same 351 symbolic rule tasks after every token is remapped to unfamiliar symbols (P1), the solver states its inferred rule and reasoning, then its answer at near-zero temperature (1e-7) is judged against the code-executed ground truth by a GPT-5-mini judge; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro | 34.8 |
| 2 | Claude Sonnet 4.5 | 32.5 |
| 3 | GPT-5.2 | 28.5 |
| 4 | GPT-5 | 23.6 |
| 5 | Gemini 2.5 Flash | 23.4 |
| 6 | GPT-5 Mini | 22.8 |
| 7 | Qwen 3 32B | 21.9 |
| 8 | GLM-4.6 | 21.4 |
| 9 | O4 Mini | 20.2 |
| 10 | DeepSeek V3.2 | 18.2 |
| 11 | Qwen 3 14B | 17.1 |
| 12 | GPT-4o Mini | 15.1 |
| 13 | Gemini 3 Flash | 10.3 |
Interactive version: theaggregate.ai/benchmark?slug=a2rbench-symbolic-tasks-remapped-symbols · How It Works · Data refreshed daily, snapshot 2026-10-07.