FAM-Bench - Comparative Ranking (CoT and Knowledge Injection): leaderboard

Metric: Top-1 accuracy (%) on Task 2: rank four candidate dishes for a health condition and match the gold most suitable dish, CoT and Knowledge Injection prompting mode (reasoning steps and injected dietary rules combined), 1,000 comparative queries over four candidate dishes with recipe text and image, 13 health conditions, nutrition-expert verified, greedy decoding with JSON output; chance 25; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 5 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.641.5
2Gemini 2.5 Pro41.3
3GPT-5.437.4
4Gemma 3 12B (IT)32.3
5Qwen 3 VL 8B Instruct30.8

Interactive version: theaggregate.ai/benchmark?slug=fam-bench-comparative-ranking-cot-and-knowledge-injection · How It Works · Data refreshed daily, snapshot 2026-10-07.