VoxelCodeBench - Symbolic Reasoning: leaderboard

Metric: Shape correctness (%) on the symbolic interpretation (80 tasks: coordinate-based primitive placement and pattern reconstruction) tasks, judged by a human annotator; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 4 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.590.3
2GPT-587.5
3Claude Opus 475
4Gemini 3 Pro15.2

Interactive version: theaggregate.ai/benchmark?slug=voxelcodebench-symbolic-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.