flan-t5-xl: benchmark results
Provider: Google. Released 2022-10-21. Access: API.
Unified ELO 1398 ± 1, rank #1143 of 1392 rated models, from 13 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ConvRe - Re2Text Easy | 91.5 | Accuracy (%) | 88.9 |
| InstructEval - Problem Solving | 47.4 | Average (%, MMLU/BBH/DROP/CRASS/HumanEval) | 87.8 |
| InstructEval - Alignment (HHH) | 75 | Average (%, harmless/helpful/honest) | 84 |
| ConvRe - Text2Re Easy | 90.6 | Accuracy (%) | 77.8 |
| Open LLM Leaderboard - MuSR | 11.85 | Score | 63.8 |
| ConvRe | 52 | Average Score (%) | 44.4 |
| Open LLM Leaderboard - BBH | 22.84 | Score | 34.3 |
| Open LLM Leaderboard - MMLU-Pro | 12.74 | Score | 22.5 |
| ConvRe - Text2Re Hard | 17.8 | Accuracy (%) | 22.2 |
| Open LLM Leaderboard - IFEval | 22.37 | Score | 15.4 |
| Open LLM Leaderboard - GPQA | 0.34 | Score | 6.3 |
| Open LLM Leaderboard - MATH Level 5 | 0.76 | Score | 5.6 |
Interactive version: theaggregate.ai/model?slug=flan-t5-xl · How It Works · Data refreshed daily, snapshot 2026-09-05.