flan-t5-xxl: benchmark results

Provider: Google. Released 2022-10-21. Access: API.

Unified ELO 1418 ± 1, rank #1078 of 1392 rated models, from 16 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
ConvRe - Text2Re Easy96.8Accuracy (%)100
InstructEval - Problem Solving50.8Average (%, MMLU/BBH/DROP/CRASS/HumanEval)95.9
InstructEval - Alignment (HHH)76.4Average (%, harmless/helpful/honest)88
CMMMU (Validation)36.8Validation Overall (%)79.4
CMMMU31.2Test Overall (%)58.8
CyberMetric72.38Accuracy (%)58.3
Open LLM Leaderboard - MuSR11.19Score58
Open LLM Leaderboard - BBH30.12Score52.2
ConvRe - Re2Text Easy79.4Accuracy (%)33.3
ConvRe - Re2Text Hard20.7Accuracy (%)33.3
Open LLM Leaderboard - GPQA2.68Score24.8
Open LLM Leaderboard - MMLU-Pro14.92Score24

Interactive version: theaggregate.ai/model?slug=flan-t5-xxl · How It Works · Data refreshed daily, snapshot 2026-09-05.