LiveBench Paraphrase — leaderboard

Metric: Score. Source: livebench.ai. 73 models tracked.

Top models

#ModelScore
1Llama 3.1 70B Instruct82.25
2Llama 3.1 405B Instruct81.13
3Gemini 1.5 Pro (Preview 0801)78.63
4Gemini 1.5 Pro (Preview 0827)77.58
5Gemini 1.5 Flash (Preview 0827)73.58
6Claude 3.5 Sonnet (20240620)72.22
7Gemma 2 27B (IT)71.13
8GPT-4o (2024-05-13)70.8
9Command-R+70.22
10GPT-4o (2024-08-06)70.08
11Claude 3 Opus (20240229)69.38
12Gemini 1.5 Pro (0514)69.27
13c4ai-command-r-08-202469.05
14Qwen 2 72B Instruct68.73
15DeepSeek V2.568.72

Interactive version: theaggregate.ai/benchmark?slug=livebench-paraphrase · How the rankings work · Data refreshed daily, snapshot 2026-07-22.