WANDR: leaderboard
Metric: Soft F1 (%; x100 of the 0-1 harmonic mean of soft precision and recall, where incomplete members earn partial credit and recall zero-pads unmet required counts; one run per system over all 500 WANDR tasks, each asking for n x m x k qualifying records with cited source pages; a task-specific judge (pinned GPT-5.4 configuration) re-fetches every cited page and verifies each record's claims (verdict_full); per-task scores averaged without weighting, errored trials counted as 0). Source: arxiv.org. Saturation forecast: Around 2029. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.8 (High) | 24.9 |
| 2 | GPT-5.5 (High) | 12.1 |
Interactive version: theaggregate.ai/benchmark?slug=wandr · How It Works · Data refreshed daily, snapshot 2026-09-26.