AutoLogi — leaderboard

AutoLogi evaluates model capability on intelligence & reasoning tasks from the linked upstream source with Score as the primary reported metric.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 8 models tracked.

Top models

#ModelScore
1GPT-4o (2024-08-06)62.04
2Claude 3.5 Sonnet60.97
3Llama 3.1 405B Instruct59.42
4Llama 3.1 70B Instruct53.79
5Qwen 2.5 72B Instruct53.5
6Qwen 2.5 7B Instruct27.28
7Llama 3.1 8B Instruct25.53
8GPT-3.5 Turbo18.93

Interactive version: theaggregate.ai/benchmark?slug=autologi · How the rankings work · Data refreshed daily, snapshot 2026-07-22.