WANDR (Hard F1): leaderboard
Metric: Hard F1 (%; x100 of the 0-1 harmonic mean of hard precision and recall, credit only for submitted members whose required subtree is fully correct, recall zero-padded to the required count; one run per system over all 500 WANDR tasks, each asking for n x m x k qualifying records with cited source pages; a task-specific judge (pinned GPT-5.4 configuration) re-fetches every cited page and verifies each record's claims (verdict_full); per-task scores averaged without weighting, errored trials counted as 0). Source: arxiv.org. Saturation forecast: Around 2031. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.8 (High) | 7.2 |
| 2 | GPT-5.5 (High) | 3.5 |
Interactive version: theaggregate.ai/benchmark?slug=wandr-hard-f1 · How It Works · Data refreshed daily, snapshot 2026-09-26.