MADE: leaderboard

Metric: Macro-F1 (%, the 0-1 score times 100) over all 1,154 labels with ten-shot prompting with kNN-retrieved training reports, the task instructions and the label list, on the 10,288-report stratified (truncated) MADE test set of FDA medical device adverse event reports from July 2024 to June 2025 with 1,154 hierarchical IMDRF product and patient problem labels, greedy decoding (GPT-5 at its default medium reasoning effort); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1GPT-554
2Qwen 3 235B A22B (Thinking)49
3DeepSeek R1 052848
4Qwen 3 30B A3B (Thinking)45
5Qwen 3 235B A22B Instruct44
6GPT-4.143
7GLM 4.5 Air42
8GPT-OSS-120B40
9Qwen 3 4B (Reasoning)38
10Llama 3.1 70B Instruct30
11Qwen 3 4B Instruct29
12Qwen 3 30B A3B Instruct22
13Kimi K29
14Llama 3.1 8B Instruct8

Interactive version: theaggregate.ai/benchmark?slug=made · How It Works · Data refreshed daily, snapshot 2026-10-07.