WANDR: leaderboard

Metric: Soft F1 (%; x100 of the 0-1 harmonic mean of soft precision and recall, where incomplete members earn partial credit and recall zero-pads unmet required counts; one run per system over all 500 WANDR tasks, each asking for n x m x k qualifying records with cited source pages; a task-specific judge (pinned GPT-5.4 configuration) re-fetches every cited page and verifies each record's claims (verdict_full); per-task scores averaged without weighting, errored trials counted as 0). Source: arxiv.org. Saturation forecast: Around 2029. 6 models tracked.

Top models

#ModelScore
1Claude Opus 4.8 (High)24.9
2GPT-5.5 (High)12.1

Interactive version: theaggregate.ai/benchmark?slug=wandr · How It Works · Data refreshed daily, snapshot 2026-09-26.