LiveBench Paraphrase — leaderboard
Metric: Score. Source: livebench.ai. 73 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Llama 3.1 70B Instruct | 82.25 |
| 2 | Llama 3.1 405B Instruct | 81.13 |
| 3 | Gemini 1.5 Pro (Preview 0801) | 78.63 |
| 4 | Gemini 1.5 Pro (Preview 0827) | 77.58 |
| 5 | Gemini 1.5 Flash (Preview 0827) | 73.58 |
| 6 | Claude 3.5 Sonnet (20240620) | 72.22 |
| 7 | Gemma 2 27B (IT) | 71.13 |
| 8 | GPT-4o (2024-05-13) | 70.8 |
| 9 | Command-R+ | 70.22 |
| 10 | GPT-4o (2024-08-06) | 70.08 |
| 11 | Claude 3 Opus (20240229) | 69.38 |
| 12 | Gemini 1.5 Pro (0514) | 69.27 |
| 13 | c4ai-command-r-08-2024 | 69.05 |
| 14 | Qwen 2 72B Instruct | 68.73 |
| 15 | DeepSeek V2.5 | 68.72 |
Interactive version: theaggregate.ai/benchmark?slug=livebench-paraphrase · How the rankings work · Data refreshed daily, snapshot 2026-07-22.