PSEBench - Boundary Condition Hit Rate: leaderboard

Metric: Boundary condition hit rate (%, M4): share of the case boundary conditions that the rationale invokes consistently, judged by GPT-5.2, pooled over correctly triaged cases, on the 5,074-case MN29 benchmark (3,455 complete, 1,362 missing-information and 257 uncertain cases built from clause cards of the Minnesota 29 reportable adverse health events), agentic environment where the model may ASK a GPT-5.2 information provider within 10 turns and then answers with a verdict, clause, legal evidence and rationale; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1GPT-5.588
2Claude Sonnet 4.685.8
3GPT-583.9
4Qwen 3 235B A22B 2507 Instruct81.6
5Gemini 3.1 Pro (Preview)78.6
6Claude Opus 4.778.5
7Gemini 2.5 Flash77.5
8DeepSeek R176.5
9GPT-5 Nano72.5
10Mistral Small 3.271.4
11GPT-OSS-120B68.5
12Llama 3.1 8B Instruct62.1

Interactive version: theaggregate.ai/benchmark?slug=psebench-boundary-condition-hit-rate · How It Works · Data refreshed daily, snapshot 2026-09-29.