HumanLikeness - Discourse-2 — leaderboard

Metric: Humanlike Score (%). Source: huggingface.co. 20 models tracked.

Top models

#ModelScore
1Llama 2 13B Chat Base78.85
2Phi-3 Mini 4K Instruct76.18
3zephyr-7B-alpha74.58
4Mistral Nemo Instruct (2407)67.84
5Llama 3.1 70B Instruct63.12
6GPT-4o62.11
7starchat2-15B-v0.161.81
8CodeLlama-34B-Instruct-hf60.74
9zephyr-7B-beta59.56
10GPT-4o Mini57.32
11Llama 3.1 8B Instruct55.57
12Llama 3 8B Instruct54.28
13Mixtral 8x7B Instruct (v0.1)52.18
14c4ai-command-r-plus50.44
15Llama 3 70B Instruct50.4

Interactive version: theaggregate.ai/benchmark?slug=humanlikeness-discourse-2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.