LMArena Preference Proxy — leaderboard

Preference proxy evaluation from LMArena, measuring how well LLM evaluators reproduce preference signals across math, code, instruction-following, and knowledge tasks.

Metric: Evaluator Accuracy (%). Source: huggingface.co. 4 models tracked.

Top models

#ModelScore
1Gemma 2 9B (IT)64.63
2Llama 3 8B Instruct64.01
3GPT-4o Mini (2024-07-18)58.83
4Claude 3 Haiku (20240307)57.39

Interactive version: theaggregate.ai/benchmark?slug=lmarena-preference-proxy · How the rankings work · Data refreshed daily, snapshot 2026-07-22.