MADE - Head Labels: leaderboard
Metric: Macro-F1 (%, the 0-1 score times 100) over the 144 head labels (more than 1 percent of training reports) with ten-shot prompting with kNN-retrieved training reports, the task instructions and the label list, on the 10,288-report stratified (truncated) MADE test set of FDA medical device adverse event reports from July 2024 to June 2025 with 1,154 hierarchical IMDRF product and patient problem labels, greedy decoding (GPT-5 at its default medium reasoning effort); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 68 |
| 2 | DeepSeek R1 0528 | 62 |
| 3 | Qwen 3 235B A22B (Thinking) | 62 |
| 4 | Qwen 3 235B A22B Instruct | 60 |
| 5 | GPT-4.1 | 59 |
| 6 | Qwen 3 30B A3B (Thinking) | 58 |
| 7 | GPT-OSS-120B | 57 |
| 8 | GLM 4.5 Air | 56 |
| 9 | Qwen 3 4B (Reasoning) | 53 |
| 10 | Llama 3.1 70B Instruct | 50 |
| 11 | Qwen 3 4B Instruct | 49 |
| 12 | Qwen 3 30B A3B Instruct | 48 |
| 13 | Llama 3.1 8B Instruct | 28 |
| 14 | Kimi K2 | 18 |
Interactive version: theaggregate.ai/benchmark?slug=made-head-labels · How It Works · Data refreshed daily, snapshot 2026-10-07.