LifeAgentBench - Context Prompting: leaderboard

Metric: Answer Accuracy (%; single-user questions, pre-filtered records in the prompt, 32 new tokens). Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1DeepSeek V4 Pro58.89
2GPT-4o57.02
3Qwen 3.5 9B45.55
4Gemini 2.5 Flash Lite44.81
5Claude 3 Haiku35.3
6Gemma 2 9B (IT)24.44
7Llama 3.1 8B Instruct20.65
8Phi-3.5-mini-instruct20.57
9Llama 3.2 3B Instruct20.18

Interactive version: theaggregate.ai/benchmark?slug=lifeagentbench-context-prompting · How It Works · Data refreshed daily, snapshot 2026-09-25.