PyraMathBench - Sorting: leaderboard
Metric: Score (0-100) on the sorting subtask (582 items: order numbers or objects), score (0-100) per item: numeric answers score 100 within 1e-4 and fall linearly to 0 at 1 percent relative error, expressions are matched with Math-Verify, text answers by MathBERT cosine similarity of at least 0.9 and choices by exact match; zero-shot chain of thought, median of three samples at temperature 0.7; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek R1 | 97.4 |
| 2 | GPT-4o Mini (2024-07-18) | 96.6 |
| 3 | GPT-4o (2024-11-20) | 95.8 |
| 4 | Claude Opus 4.6 | 94.7 |
Interactive version: theaggregate.ai/benchmark?slug=pyramathbench-sorting · How It Works · Data refreshed daily, snapshot 2026-09-29.