LifeAgentBench - Database-Augmented Prompting: leaderboard
Metric: Answer Accuracy (%; all questions, model-written read-only SQL retrieval then answer). Source: arxiv.org. Saturation forecast: Around March 2027. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V4 Pro | 55.67 |
| 2 | Gemini 2.5 Flash Lite | 39.04 |
| 3 | GPT-4o | 34.71 |
| 4 | Claude 3 Haiku | 29.3 |
| 5 | Qwen 3.5 9B | 25.76 |
| 6 | Llama 3.1 8B Instruct | 21.53 |
| 7 | Phi-3.5-mini-instruct | 16.16 |
| 8 | Gemma 2 9B (IT) | 14.54 |
| 9 | Llama 3.2 3B Instruct | 13.47 |
Interactive version: theaggregate.ai/benchmark?slug=lifeagentbench-database-augmented-prompting · How It Works · Data refreshed daily, snapshot 2026-09-25.