Gemini 2.5 Flash (Preview 09-2025) (Thinking) — benchmark results
September 2025 Gemini 2.5 Flash preview snapshot evaluated with thinking enabled. Provider: Google. Released 2025-09-25. Access: API.
Unified ELO 1690 ± 15, rank #293 of 1776 rated models, from 19 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| UGI Leaderboard | 50.35 | UGI Score | 90.2 |
| UGI - Natural Intelligence | 43.92 | NatInt Score | 88.7 |
| UGI - Writing | 43.88 | Writing Score | 84.7 |
| Vals AI LegalBench | 82.62 | Accuracy (%) | 62.4 |
| Vals AI MMMU | 80.75 | Accuracy (%) | 61 |
| Vals AI TaxEval v2 | 72.4 | Accuracy (%) | 59.4 |
| Vals AI MedQA | 91.17 | Accuracy (%) | 57.4 |
| Vals AI MMLU-Pro | 83.66 | Accuracy (%) | 53.7 |
| Vals AI MedScribe | 78.5 | Accuracy (%) | 51.4 |
| Vals AI LiveCodeBench | 76.21 | Accuracy (%) | 50.4 |
| Vals AI GPQA | 76.52 | Accuracy (%) | 48.4 |
| Vals AI CorpFin v2 | 59.75 | Accuracy (%) | 47.5 |
Interactive version: theaggregate.ai/model?slug=gemini-2-5-flash-preview-09-2025-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.