AlpacaEval: leaderboard
Automatic instruction-following evaluator comparing model responses against a reference using GPT-4 judgments and length-controlled win rates.
Source: tatsu-lab.github.io. Status: saturated.
Interactive version: theaggregate.ai/benchmark?slug=alpacaeval · How It Works · Data refreshed daily, snapshot 2026-09-05.