flan-t5-xxl — benchmark results
Provider: Google. Released 2022-10-21. Access: API.
Unified ELO 1336 ± 43, rank #1454 of 1776 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ConvRe - Text2Re Easy | 96.8 | Accuracy (%) | 100 |
| CMMMU (Validation) | 36.8 | Validation Overall (%) | 79.4 |
| CMMMU | 31.2 | Test Overall (%) | 58.8 |
| CyberMetric | 72.38 | Accuracy (%) | 58.3 |
| Open LLM Leaderboard - MuSR | 11.19 | Score | 58 |
| Open LLM Leaderboard - BBH | 30.12 | Score | 52.2 |
| ConvRe - Re2Text Easy | 79.4 | Accuracy (%) | 33.3 |
| ConvRe - Re2Text Hard | 20.7 | Accuracy (%) | 33.3 |
| Open LLM Leaderboard - GPQA | 2.68 | Score | 24.8 |
| Open LLM Leaderboard - MMLU-Pro | 14.92 | Score | 24 |
| Open LLM Leaderboard - IFEval | 22 | Score | 15 |
| ConvRe | 50.4 | Average Score (%) | 11.1 |
Interactive version: theaggregate.ai/model?slug=flan-t5-xxl · How the rankings work · Data refreshed daily, snapshot 2026-07-22.