OMLAB Open Agent Math Leaderboard — leaderboard
Open Agent Leaderboard math-reasoning track comparing prompting and agent algorithms across GSM8K, AQuA, and MATH-500 with score and cost metrics.
Metric: Average Score (%). Source: huggingface.co. 66 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen2.5-72B-Instruct (SC-CoT) | 86.67 |
| 2 | Qwen2.5-72B-Instruct (CoT) | 86.43 |
| 3 | gpt-4o (SC-CoT) | 85.07 |
| 4 | Llama-3.3-70B-Instruct (SC-CoT) | 84.09 |
| 5 | Llama-3.3-70B-Instruct (CoT) | 82.86 |
| 6 | gpt-4o (CoT) | 81.59 |
| 7 | Llama-3.3-70B-Instruct (IO) | 81.45 |
| 8 | Qwen2.5-7B-Instruct (SC-CoT) | 80.57 |
| 9 | Qwen2.5-72B-Instruct (IO) | 80.34 |
| 10 | Qwen2.5-7B-Instruct (CoT) | 78.73 |
Interactive version: theaggregate.ai/benchmark?slug=omlab-open-agent-math-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.