PyraMathBench - Sorting: leaderboard

Metric: Score (0-100) on the sorting subtask (582 items: order numbers or objects), score (0-100) per item: numeric answers score 100 within 1e-4 and fall linearly to 0 at 1 percent relative error, expressions are matched with Math-Verify, text answers by MathBERT cosine similarity of at least 0.9 and choices by exact match; zero-shot chain of thought, median of three samples at temperature 0.7; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.

Top models

#ModelScore
1DeepSeek R197.4
2GPT-4o Mini (2024-07-18)96.6
3GPT-4o (2024-11-20)95.8
4Claude Opus 4.694.7

Interactive version: theaggregate.ai/benchmark?slug=pyramathbench-sorting · How It Works · Data refreshed daily, snapshot 2026-09-29.