HumanLikeness - Discourse-1 — leaderboard

Metric: Humanlike Score (%). Source: huggingface.co. 20 models tracked.

Top models

#ModelScore
1Phi-3 Mini 4K Instruct79.7
2Llama 3.1 70B Instruct79.69
3Llama 3.1 8B Instruct79.26
4CodeLlama-34B-Instruct-hf79.12
5Mistral Nemo Instruct (2407)78.61
6Yi 1.5 34B Chat77.91
7c4ai-command-r-plus77.7
8Llama 3 8B Instruct77.02
9zephyr-7B-alpha76.23
10GPT-3.5 Turbo76.07
11Llama 2 13B Chat Base75.97
12Llama 3 70B Instruct75.54
13Llama 2 7B Chat75.48
14GPT-4o75.43
15starchat2-15B-v0.175.11

Interactive version: theaggregate.ai/benchmark?slug=humanlikeness-discourse-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.