Gemini 2.5 Flash — benchmark results
Google's speed-oriented Gemini 2.5 Flash model, balancing capability with lower latency. Provider: Google. Released 2025-06-17. Access: API.
Unified ELO 1635 ± 8, rank #401 of 1776 rated models, from 701 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - ocw_connections | 1.19 | Dataset z-score | 100 |
| CommonWhy | 27.7 | BS-Prec (Long-Tail) (self-reported) | 100 |
| Evals for Every Language - Language wuu | 44.94 | Average Score (%) | 100 |
| Galileo Agent - Telecom TSQ | 95 | Task Success Quality (%) | 100 |
| K-FinHallu | 86.4 | Accuracy (self-reported) | 100 |
| Ko-AgentBench - L6 Efficient Tool Utilization | 33.33 | Efficiency Score (%) | 100 |
| MedFrameQA | 54.75 | Average Accuracy (self-reported) | 100 |
| Medical MM Leaderboard - MMMU-Med | 76.9 | Accuracy (%) | 100 |
| Medical MM Leaderboard - MedXQA | 52.8 | Accuracy (%) | 100 |
| OmniToM | 85.95 | Stage 2 Overall Acc (self-reported) | 100 |
| Open Portuguese LLM - ASSIN2 STS | 87.15 | Pearson Correlation (×100) | 100 |
| Physical Reasoning - CausalVQA | 61.66 | Accuracy (%) | 100 |
Interactive version: theaggregate.ai/model?slug=gemini-2-5-flash · How the rankings work · Data refreshed daily, snapshot 2026-07-22.