Gemini 3.1 Flash Lite (Non-reasoning): benchmark results
Provider: Google. Released 2026-05-07. Access: API.
Unified ELO 1585 ± 23, rank #708 of 2055 rated models, from 18 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| UAV-DualCog - Landmark Visibility Counting | 59.9 | Visibility-count accuracy (%; 1,024 flight videos in which t | 96.4 |
| UAV-DualCog - Self-Relative Position | 33.7 | Answer accuracy (%; 1,024 questions asking where the UAV is | 80 |
| UAV-DualCog - Flight Behavior Recognition (Composite) | 22.8 | Composite-level behavior accuracy (%; 1,024 first-person UAV | 65.4 |
| UAV-DualCog - Future Observation Prediction | 30.1 | Answer accuracy (%; 1,024 questions asking which view the UA | 61.8 |
| UAV-DualCog - Landmark-Relative Direction | 47.6 | Answer accuracy (%; 1,024 questions asking where a landmark | 55.7 |
| Hy-MultiTurn - Action Suppression | 56.2 | Importance-weighted score (%) on the 35 action suppression d | 47.6 |
| Hy-MultiTurn - Constraint Memory | 34.4 | Importance-weighted score (%) on the 34 constraint memory di | 28.6 |
| UAV-DualCog - Landmark-Driven Action Decision | 35.9 | Answer accuracy (%; 1,024 questions asking which way the UAV | 20 |
| UAV-DualCog - Flight Behavior Recognition (Atomic) | 23.3 | Atomic-level behavior accuracy (%; the atomic flight actions | 19.2 |
| LLM Milgram Obedience - Baseline Full-Obedience Rate | 100 | Valid baseline sessions ending at 450 V (%; lower is less ob | 5.3 |
| LLM Milgram Obedience - Baseline Mean Breakoff Voltage | 450 | Mean last shock administered in valid baseline sessions (vol | 5.3 |
| Hy-MultiTurn | 38.9 | Importance-weighted score (%): share of the importance-weigh | 0 |
Interactive version: theaggregate.ai/model?slug=gemini-3-1-flash-lite-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-29.