DFIR ATT&CK Technique Classification (3-Shot CoT): leaderboard
Metric: Micro-F1 (%; multi-label ATT&CK technique prediction for 2,076 sentences from 83 The DFIR Report incident reports, 211-technique label space; three in-context examples with chain-of-thought prompting, temperature 0; models run at Q4_K_M 4-bit quantization). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek-V2.5 [Q4_K_M] | 22 |
| 2 | Llama-3.1-70B-Instruct [Q4_K_M] | 21 |
| 3 | gpt-oss-120b [Q4_K_M] | 20 |
| 4 | gemma-3-27b-it [Q4_K_M] | 17 |
| 5 | gemma-3-12b-it [Q4_K_M] | 13 |
| 6 | Llama-3.1-8B-Instruct [Q4_K_M] | 11 |
| 7 | gpt-oss-20b [Q4_K_M] | 0 |
Interactive version: theaggregate.ai/benchmark?slug=dfir-att-ck-technique-classification-3-shot-cot · How It Works · Data refreshed daily, snapshot 2026-09-26.