flan-t5-xxl: benchmark results
Provider: Google. Released 2022-10-21. Access: API.
Unified ELO 1418 ± 1, rank #1078 of 1392 rated models, from 16 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ConvRe - Text2Re Easy | 96.8 | Accuracy (%) | 100 |
| InstructEval - Problem Solving | 50.8 | Average (%, MMLU/BBH/DROP/CRASS/HumanEval) | 95.9 |
| InstructEval - Alignment (HHH) | 76.4 | Average (%, harmless/helpful/honest) | 88 |
| CMMMU (Validation) | 36.8 | Validation Overall (%) | 79.4 |
| CMMMU | 31.2 | Test Overall (%) | 58.8 |
| CyberMetric | 72.38 | Accuracy (%) | 58.3 |
| Open LLM Leaderboard - MuSR | 11.19 | Score | 58 |
| Open LLM Leaderboard - BBH | 30.12 | Score | 52.2 |
| ConvRe - Re2Text Easy | 79.4 | Accuracy (%) | 33.3 |
| ConvRe - Re2Text Hard | 20.7 | Accuracy (%) | 33.3 |
| Open LLM Leaderboard - GPQA | 2.68 | Score | 24.8 |
| Open LLM Leaderboard - MMLU-Pro | 14.92 | Score | 24 |
Interactive version: theaggregate.ai/model?slug=flan-t5-xxl · How It Works · Data refreshed daily, snapshot 2026-09-05.