A2RBench - Symbolic Tasks (Original Symbols): leaderboard

Metric: Accuracy (%) on the 351 symbolic (structure-based) rule tasks with their original symbols (P0), the solver states its inferred rule and reasoning, then its answer at near-zero temperature (1e-7) is judged against the code-executed ground truth by a GPT-5-mini judge; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 14 models tracked.

Top models

#ModelScore
1GPT-541.3
2GPT-5 Mini40.2
3Gemini 3 Pro39.3
4O4 Mini36.5
5Claude Sonnet 4.534.2
6GPT-5.232.5
7Gemini 2.5 Flash29.3
8GLM-4.624.8
9GPT-4o Mini23.9
10DeepSeek V3.223.4
11Qwen 3 32B23.1
12Qwen 3 14B16.8
13Gemini 3 Flash15.7

Interactive version: theaggregate.ai/benchmark?slug=a2rbench-symbolic-tasks-original-symbols · How It Works · Data refreshed daily, snapshot 2026-10-07.