Gemini 2.0 Flash (01-21) (Thinking) — benchmark results
January 21 Gemini 2.0 Flash thinking experimental snapshot, kept separate from the generic thinking row. Provider: Google. Released 2025-01-21. Access: API.
Unified ELO 1630 ± 29, rank #412 of 1776 rated models, from 23 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| U-MATH - Algebra | 89.44 | Accuracy (%) | 98.5 |
| U-MATH | 83.64 | Accuracy (%) | 97 |
| U-MATH - Differential Calculus | 70.45 | Accuracy (%) | 97 |
| U-MATH - Multivariable Calculus | 83.71 | Accuracy (%) | 97 |
| U-MATH - Integral Calculus | 82.21 | Accuracy (%) | 95.5 |
| U-MATH - Sequences & Series | 88.31 | Accuracy (%) | 90.9 |
| U-MATH - Precalculus | 92.5 | Accuracy (%) | 87.9 |
| μ-MATH | 81.2 | Accuracy (%) | 76.5 |
| NYT Connections Original | 37 | Score (%) | 73.3 |
| Generalization V1 (Lechmazur) | 1.84 | Avg Rank (lower is better) | 68.8 |
| Confabulation Leaderboard (Lechmazur) | 14.85 | Confabulation rate % (lower is better) | 65.1 |
| Wolfram LLM Benchmarking Project | 46 | Correct Functionality (%) | 61.8 |
Interactive version: theaggregate.ai/model?slug=gemini-2-0-flash-01-21-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.