Gemini 2.5 Flash (Preview 04-17) (Thinking) — benchmark results
April 17 Gemini 2.5 Flash preview snapshot evaluated with thinking enabled. Provider: Google. Released 2025-04-17. Access: API.
Unified ELO 1675 ± 31, rank #312 of 1776 rated models, from 10 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Arena RU | 1159 | Arena Elo | 96.7 |
| SpeechMap Compliance | 75.6 | % Requests Completed | 68.8 |
| Vals AI MATH 500 | 91.8 | Accuracy (%) | 67 |
| ReliableMath - Prudence | 0.1 | Score (%) | 57.9 |
| Vals AI MedQA | 91.02 | Accuracy (%) | 55.3 |
| Vals AI TaxEval v2 | 70.52 | Accuracy (%) | 40.6 |
| ReliableMath - Precision | 50.8 | Score (%) | 36.8 |
| Vals AI MMMU | 71.92 | Accuracy (%) | 32.9 |
| Vals AI MortgageTax | 57.51 | Accuracy (%) | 26.7 |
| Vals AI LiveCodeBench | 46.87 | Accuracy (%) | 18.1 |
Interactive version: theaggregate.ai/model?slug=gemini-2-5-flash-preview-04-17-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.