MADE - Jaccard: leaderboard

Metric: Sample-level Jaccard index (%, the 0-1 score times 100) between the predicted and gold label sets with ten-shot prompting with kNN-retrieved training reports, the task instructions and the label list, on the 10,288-report stratified (truncated) MADE test set of FDA medical device adverse event reports from July 2024 to June 2025 with 1,154 hierarchical IMDRF product and patient problem labels, greedy decoding (GPT-5 at its default medium reasoning effort); higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 15 models tracked.

Top models

#ModelScore
1GPT-557
2GPT-4.157
3DeepSeek R1 052850
4Qwen 3 235B A22B Instruct49
5Qwen 3 235B A22B (Thinking)48
6Qwen 3 30B A3B (Thinking)47
7GPT-OSS-120B45
8GLM 4.5 Air44
9Llama 3.1 70B Instruct43
10Qwen 3 4B (Reasoning)43
11Qwen 3 30B A3B Instruct43
12Qwen 3 4B Instruct41
13Llama 3.1 8B Instruct22
14Kimi K27

Interactive version: theaggregate.ai/benchmark?slug=made-jaccard · How It Works · Data refreshed daily, snapshot 2026-10-07.