FoodGuardBench: leaderboard

Metric: Attack success rate (%) of the AutoDAN and PAP jailbreak versions (4,464 prompts; AutoDAN prompts optimized on Llama 3.1 8B and transferred) of the harmful food-safety queries of FoodGuardBench (2,339 vanilla queries grounded in FDA food-code principles across eight hazard categories), temperature 0.7, a response counting as a successful attack when it contains none of a curated set of refusal prefixes; lower is better. Source: arxiv.org. 9 models tracked.

Top models

#ModelScore
1Claude 3.7 Sonnet34.59
2GPT-4.150.54
3GPT-4o54.84
4Llama 3.3 70B Instruct55.49
5Qwen 3 8B (Non-reasoning)56.38
6GLM-4 32B60.84
7Qwen 3 32B (Non-reasoning)62.76
8Qwen 2.5 7B62.99
9Mistral Small 467.58

Interactive version: theaggregate.ai/benchmark?slug=foodguardbench · How It Works · Data refreshed daily, snapshot 2026-10-07.