HumanLikeness - Overall — leaderboard

Psycholinguistic humanlikeness benchmark comparing model response distributions with 2,000 human participants across sound, word, syntax, meaning, and discourse tasks.

Metric: Overall Humanlike (%). Source: huggingface.co. Status: saturation imminent. 20 models tracked.

Top models

#ModelScore
1Llama 3.1 70B Instruct66.91
2Llama 3.1 8B Instruct66.2
3Phi-3 Mini 4K Instruct64.73
4Mistral Nemo Instruct (2407)63.97
5Llama 2 13B Chat Base63.26
6CodeLlama-34B-Instruct-hf62.16
7c4ai-command-r-plus61.09
8Llama 3 8B Instruct61.01
9starchat2-15B-v0.160.84
10GPT-4o58.98
11GPT-3.5 Turbo58.66
12Yi 1.5 34B Chat58.5
13Llama 2 7B Chat57.6
14Llama 3 70B Instruct57.14
15zephyr-7B-alpha56.89

Interactive version: theaggregate.ai/benchmark?slug=humanlikeness-overall · How the rankings work · Data refreshed daily, snapshot 2026-07-22.