OMLAB Open Agent Math Leaderboard — leaderboard

Open Agent Leaderboard math-reasoning track comparing prompting and agent algorithms across GSM8K, AQuA, and MATH-500 with score and cost metrics.

Metric: Average Score (%). Source: huggingface.co. 66 models tracked.

Top models

#ModelScore
1Qwen2.5-72B-Instruct (SC-CoT)86.67
2Qwen2.5-72B-Instruct (CoT)86.43
3gpt-4o (SC-CoT)85.07
4Llama-3.3-70B-Instruct (SC-CoT)84.09
5Llama-3.3-70B-Instruct (CoT)82.86
6gpt-4o (CoT)81.59
7Llama-3.3-70B-Instruct (IO)81.45
8Qwen2.5-7B-Instruct (SC-CoT)80.57
9Qwen2.5-72B-Instruct (IO)80.34
10Qwen2.5-7B-Instruct (CoT)78.73

Interactive version: theaggregate.ai/benchmark?slug=omlab-open-agent-math-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.