PaLM-2 (Bison) — benchmark results

Google's mid-size PaLM 2 tier behind Vertex AI's chat-bison API, announced at I/O in May 2023 and since retired. Provider: Google. Released 2023-05-10. Access: API.

Unified ELO 1495 ± 22, rank #820 of 1776 rated models, from 10 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM NaturalQuestions (Open)81.26F1 (%)97.8
HELM WMT 201424.08BLEU-4 (%)97.8
HELM v2 Lite - LegalBench62.96Exact Match (%)85
HELM v2 Lite - HumanEval (Code)29Pass@1 (%)76.5
HELM Lite56.77Mean win rate (self-reported)63.6
HELM NaturalQuestions (Closed)39F1 (%)57.8
HELM v2 Lite - MedQA67.31Exact Match (%)57.5
HELM (Stanford)52.64Mean Win Rate (%)56.7
HELM NarrativeQA71.8F1 (%)36.7
LLM-AggreFact66.13Balanced Accuracy (%)23.7

Interactive version: theaggregate.ai/model?slug=palm-2-bison · How the rankings work · Data refreshed daily, snapshot 2026-07-22.