Gemini 3.5 Flash (Low): benchmark results
Provider: Google. Released 2026-05-19. Access: API.
Unified ELO 1685 ± 1, rank #184 of 3078 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Chess Puzzles (Epoch AI) | 45 | Accuracy (%) | 93.2 |
| ObviousBench | 99.31 | Answer pass³ (%) | 92.6 |
| Epoch AI - GPQA Diamond | 88.89 | Accuracy (%) | 85.1 |
| OTIS Mock AIME 2024-25 | 88.89 | Accuracy (%) | 76.8 |
| Epoch AI - Mystery Game Puzzles | 28 | Score | 76.6 |
| GRIPS-hard | 65.3 | Accuracy (%) | 73.1 |
| FinLifeBench - Financial State - Checkpoint State Accuracy | 75.4 | Cell-level state accuracy (%) | 70 |
| GRIPS | 91.6 | Accuracy (%) | 69.2 |
| FinLifeBench - Financial State - Evidence Recall | 77.3 | Recall of gold evidence sessions (%) | 40 |
| FinLifeBench - Financial State - Granular Change Accuracy | 41.2 | GCA@15 (%) | 40 |
| FinLifeBench - Life-Event History - Event-Anchor F1 | 48.8 | F1 (%; event type with first-establishing session) | 30 |
| FinLifeBench - Life-Event History - Exact History Match | 5.3 | Checkpoints with the exact history (%) | 30 |
Interactive version: theaggregate.ai/model?slug=gemini-3-5-flash-low · How It Works · Data refreshed daily, snapshot 2026-09-19.