A2RBench - Semantic Tasks: leaderboard

Metric: Accuracy (%) on the 352 semantic (knowledge-dependent) rule tasks, the solver states its inferred rule and reasoning, then its answer at near-zero temperature (1e-7) is judged against the code-executed ground truth by a GPT-5-mini judge; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.

Top models

#ModelScore
1GPT-552
2Gemini 3 Pro48.6
3GPT-5 Mini47.7
4O4 Mini44.3
5Gemini 2.5 Flash30.4
6Claude Sonnet 4.525.3
7GPT-5.223.9
8DeepSeek V3.220.7
9GLM-4.616.8
10Gemini 3 Flash13.4
11Qwen 3 32B12.8
12GPT-4o Mini9.4
13Qwen 3 14B8.8

Interactive version: theaggregate.ai/benchmark?slug=a2rbench-semantic-tasks · How It Works · Data refreshed daily, snapshot 2026-10-07.