SAGE — leaderboard

SAGE evaluates model capability on education tasks from the linked upstream source with Score as the primary reported metric.

Metric: OA (%) (self-reported). Source: benchmarklist.com. Status: saturation imminent. 12 models tracked.

Top models

#ModelScore
1GPT-4o74.17
2Qwen 3.5 35B A3B71.79
3Llama 3.1 70B71.54
4Qwen 3.5 9B66.21
5Gemma 3 27B66.17
6Llama 3.1 8B64.71
7Gemma 3 12B64.5
8Gemma 3 4B60.79
9Qwen 3.5 4B60.5
10Llama 3.2 3B59.54
11Qwen 3.5 2B54.12
12Llama 3.2 1B53.67

Interactive version: theaggregate.ai/benchmark?slug=sage · How the rankings work · Data refreshed daily, snapshot 2026-07-22.