Gemini 2.5 Flash (Preview 04-17) (Thinking): benchmark results
April 17 Gemini 2.5 Flash preview snapshot evaluated with thinking enabled. Provider: Google. Released 2025-04-17. Access: API.
Unified ELO 1577 ± 1, rank #526 of 1761 rated models, from 10 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Arena RU | 1159 | Arena Elo | 96.1 |
| Vals AI MATH 500 | 91.8 | Accuracy (%) | 72.5 |
| SpeechMap Compliance | 75.6 | % Requests Completed | 70.7 |
| ReliableMath - Prudence | 0.1 | Score (%) | 57.9 |
| Vals AI MedQA | 91.02 | Accuracy (%) | 55.3 |
| ReliableMath - Precision | 50.8 | Score (%) | 36.8 |
| Vals AI TaxEval v2 | 70.52 | Accuracy (%) | 36.1 |
| Vals AI MMMU | 71.92 | Accuracy (%) | 29.3 |
| Vals AI MortgageTax | 57.51 | Accuracy (%) | 23.7 |
| Vals AI LiveCodeBench | 46.87 | Accuracy (%) | 16.2 |
Interactive version: theaggregate.ai/model?slug=gemini-2-5-flash-preview-04-17-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.