gemma-2-9B-it-DPO — benchmark results

Princeton NLP's DPO fine-tune of Gemma 2 9B IT on ultrafeedback-armorm, the DPO baseline counterpart to their SimPO model. Provider: Google. Released 2024-07-16. Access: API.

Unified ELO 1478 ± 72, rank #884 of 1776 rated models, from 10 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AlpacaEval 2.067.66LC Win Rate (%)95
WildBench53.22WB Score Task-Macro85.5
BenchBench81.01Aggregate Score (%)85.3
Open LLM Leaderboard - GPQA11.41Score82.7
Open LLM Leaderboard - BBH41.59Score80.8
Open LLM Leaderboard - MMLU-Pro30.26Score61.2
Open Korean LLM Leaderboard44.89Average Score (%)46.1
Open LLM Leaderboard - MATH Level 58.31Score41.9
Open LLM Leaderboard - MuSR5.65Score27.4
Open LLM Leaderboard - IFEval27.69Score24.7

Interactive version: theaggregate.ai/model?slug=gemma-2-9b-it-dpo · How the rankings work · Data refreshed daily, snapshot 2026-07-22.