LifeAgentBench - Database-Augmented Prompting: leaderboard

Metric: Answer Accuracy (%; all questions, model-written read-only SQL retrieval then answer). Source: arxiv.org. Saturation forecast: Around March 2027. 13 models tracked.

Top models

#ModelScore
1DeepSeek V4 Pro55.67
2Gemini 2.5 Flash Lite39.04
3GPT-4o34.71
4Claude 3 Haiku29.3
5Qwen 3.5 9B25.76
6Llama 3.1 8B Instruct21.53
7Phi-3.5-mini-instruct16.16
8Gemma 2 9B (IT)14.54
9Llama 3.2 3B Instruct13.47

Interactive version: theaggregate.ai/benchmark?slug=lifeagentbench-database-augmented-prompting · How It Works · Data refreshed daily, snapshot 2026-09-25.