A2RBench - Symbolic Tasks (Remapped Symbols): leaderboard

Metric: Accuracy (%) on the same 351 symbolic rule tasks after every token is remapped to unfamiliar symbols (P1), the solver states its inferred rule and reasoning, then its answer at near-zero temperature (1e-7) is judged against the code-executed ground truth by a GPT-5-mini judge; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Pro34.8
2Claude Sonnet 4.532.5
3GPT-5.228.5
4GPT-523.6
5Gemini 2.5 Flash23.4
6GPT-5 Mini22.8
7Qwen 3 32B21.9
8GLM-4.621.4
9O4 Mini20.2
10DeepSeek V3.218.2
11Qwen 3 14B17.1
12GPT-4o Mini15.1
13Gemini 3 Flash10.3

Interactive version: theaggregate.ai/benchmark?slug=a2rbench-symbolic-tasks-remapped-symbols · How It Works · Data refreshed daily, snapshot 2026-10-07.