gemma-2B-orpo: benchmark results

Provider: Google. Access: API.

Unified ELO 1328 ± 20, rank #2590 of 2928 rated models, from 12 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard v1 - GSM8K13.87Accuracy (%) (5-shot)35.8
Open LLM Leaderboard v1 - TruthfulQA MC244.53MC2 (%) (0-shot)27.5
Open LLM Leaderboard - MuSR37.28Score25.4
Open LLM Leaderboard v1 - HellaSwag73.72Normalized accuracy (%) (10-shot)24
Open LLM Leaderboard v1 - ARC Challenge49.15Normalized accuracy (%) (25-shot)23.8
Open LLM Leaderboard v1 - MMLU38.52Accuracy (%) (5-shot)20.2
Open LLM Leaderboard - IFEval24.78Score19.9
Open LLM Leaderboard - BBH34.26Score17.9
Open LLM Leaderboard v1 - WinoGrande64.33Accuracy (%) (5-shot)17.3
Open LLM Leaderboard - GPQA26.17Score15.6
Open LLM Leaderboard - MATH Level 51.89Score13.3
Open LLM Leaderboard - MMLU-Pro13.06Score10.8

Interactive version: theaggregate.ai/model?slug=gemma-2b-orpo · How It Works · Data refreshed daily, snapshot 2026-09-23.