TOFU LLaMA 10%: leaderboard
TOFU unlearning test (CMU, 2024): Llama-2-7B-chat tuned on 4,000 QA pairs about 200 fictitious authors must forget the 10% split (20 authors) and keep the rest; forget quality times model utility.
Metric: Forget Quality x Model Utility. Source: huggingface.co. 2 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | LLaMA 2-7B - Retain Model (WD=0.01) | 0.61 |
| 2 | LLaMA 2-7B - Grad. Diff. (WD=0.01) | 0 |
Interactive version: theaggregate.ai/benchmark?slug=tofu-llama-10 · How It Works · Data refreshed daily, snapshot 2026-09-05.