Linguini: leaderboard

Metric: Exact Match (%; 0 in-context examples, 894 questions). Source: arxiv.org. Saturation forecast: Estimated already saturated. 13 models tracked.

Top models

#ModelScore
1Claude 3 Opus (20240229)24.05
2GPT-4o (2024-05-13)14.65
3Claude 3 Sonnet (20240229)12.3
4GPT-4 Turbo8.72
5Llama 3 70B8.17
6GPT-4 Preview (0125)6.38
7Claude 3 Haiku (20240307)6.04
8Llama 3 70B Instruct4.81
9Llama 2 70B Base4.7
10Mixtral 8x7B (v0.1)2.46
11Qwen 1.5 110B Chat1.45
12Llama 2 70B Chat (HF)0.89
13gemma-2B0.34

Interactive version: theaggregate.ai/benchmark?slug=linguini · How It Works · Data refreshed daily, snapshot 2026-09-24.