PAL-Bench (PAL-TRACE) - Evidence Support: leaderboard

Metric: Evidence support (0-1; mean 0, 0.5 or 1 support score of the cited public photos for each matched claim; PAL-TRACE framework with the listed LLM as reconstruction backbone; 50 synthetic users, macro-averaged over users; Qwen3.6-35B-A3B semantic judge). Source: arxiv.org. Saturation forecast: Around June 2027. 9 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)0.42
2GPT-5.40.38
3DeepSeek V4 Flash0.36
4Claude Sonnet 4.50.36
5Gemini 3.1 Flash Lite0.33
6GPT-5.4 Mini0.32
7Qwen 3.6 35B A3B0.32
8Gemma 4 26B A4B (IT)0.29

Interactive version: theaggregate.ai/benchmark?slug=pal-bench-pal-trace-evidence-support · How It Works · Data refreshed daily, snapshot 2026-09-26.