FastEval: leaderboard

Aggregate LLM evaluation report combining reasoning, coding, and chat benchmarks into a total score from the FastEval project.

Metric: Total Score. Source: fasteval.github.io. Status: saturated. 33 models tracked.

Top models

#ModelScore
1GPT-4 (0613)77.78
2GPT-3.5 Turbo (0613)65.17
3GPT-3.5 Turbo (0301)63.67
4Llama 2 70B Chat48.18
5vicuna-33B-v1.343.04
6Mistral 7B Instruct (v0.1)42.66
7Llama 2 13B Chat40.94
8Llama 2 7B Chat35.36
9vicuna-7B-v1.334.5
10falcon-7B Instruct21.78

Interactive version: theaggregate.ai/benchmark?slug=fasteval · How It Works · Data refreshed daily, snapshot 2026-09-05.