HumanLikeness - Syntax-2 — leaderboard

Metric: Humanlike Score (%). Source: huggingface.co. 20 models tracked.

Top models

#ModelScore
1Llama 3 8B Instruct79.49
2Llama 3.1 8B Instruct79.21
3Llama 3.1 70B Instruct78.42
4GPT-4o78.35
5Mistral Nemo Instruct (2407)77.18
6GPT-3.5 Turbo77.07
7GPT-4o Mini76.94
8Llama 3 70B Instruct76.72
9starchat2-15B-v0.173.32
10c4ai-command-r-plus72.4
11Yi 1.5 34B Chat70.6
12Phi-3 Mini 4K Instruct69.29
13Llama 2 7B Chat67.19
14Llama 2 13B Chat Base65.18
15Mistral 7B Instruct (v0.2)58.71

Interactive version: theaggregate.ai/benchmark?slug=humanlikeness-syntax-2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.