PSEBench - Missing Slot Identification: leaderboard

Metric: Missing slot identification F1 (%, M6): micro F1 of the withheld fields the ASK queries target, on missing cases where the model asked, on the 5,074-case MN29 benchmark (3,455 complete, 1,362 missing-information and 257 uncertain cases built from clause cards of the Minnesota 29 reportable adverse health events), agentic environment where the model may ASK a GPT-5.2 information provider within 10 turns and then answers with a verdict, clause, legal evidence and rationale; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1Claude Opus 4.777.6
2Gemini 3.1 Pro (Preview)77.4
3GPT-5.577.3
4Claude Sonnet 4.676.2
5GPT-575
6Gemini 2.5 Flash71.6
7GPT-OSS-120B71.6
8Qwen 3 235B A22B 2507 Instruct69.3
9DeepSeek R166.3
10GPT-5 Nano61
11Mistral Small 3.253.5
12Llama 3.1 8B Instruct24.6

Interactive version: theaggregate.ai/benchmark?slug=psebench-missing-slot-identification · How It Works · Data refreshed daily, snapshot 2026-09-29.