HumanLikeness - Syntax-1 — leaderboard

Metric: Humanlike Score (%). Source: huggingface.co. 20 models tracked.

Top models

#ModelScore
1Phi-3 Mini 4K Instruct89.36
2starchat2-15B-v0.186.93
3zephyr-7B-alpha84.56
4Mistral Nemo Instruct (2407)84.21
5Llama 3.1 8B Instruct83.58
6CodeLlama-34B-Instruct-hf82.34
7Llama 3.1 70B Instruct81.26
8Llama 3 8B Instruct78.94
9GPT-3.5 Turbo76.33
10Llama 2 13B Chat Base74.98
11Mistral 7B Instruct (v0.3)74.89
12c4ai-command-r-plus72.36
13Mistral 7B Instruct (v0.2)72.27
14Yi 1.5 34B Chat72.24
15zephyr-7B-beta71.37

Interactive version: theaggregate.ai/benchmark?slug=humanlikeness-syntax-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.