FastEval — leaderboard

Aggregate LLM evaluation report combining reasoning, coding, and chat benchmarks into a total score from the FastEval project.

Metric: Total Score. Source: fasteval.github.io. Status: saturated. 33 models tracked.

Top models

#ModelScore
1GPT-4 (0613)77.78
2GPT-3.5 Turbo (0613)65.17
3GPT-3.5 Turbo (0301)63.67
4StableBeluga255.6
5Llama 2 70B Chat48.18
6WizardCoder-Python-34B-V1.047.15
7vicuna-33B-v1.343.04
8Mistral 7B Instruct (v0.1)42.66
9Llama 2 13B Chat40.94
10Llama 2 7B Chat35.36
11vicuna-7B-v1.334.5
12falcon-40B Instruct31.75
13mpt-7B-chat30.97
14falcon-7B Instruct21.78

Interactive version: theaggregate.ai/benchmark?slug=fasteval · How the rankings work · Data refreshed daily, snapshot 2026-07-22.