SciLitBench - Title and Abstract Screening: leaderboard
Metric: F2 score (%; recall-weighted F-score of include decisions on the frozen 1,800-record evaluation split of the manually annotated 2,000-record title and abstract seed set (52 inclusions) of a real systematic review; zero-shot prompt that states the inclusion criteria and requires explicit exclusion reasoning; Ollama, temperature 0, top-p 0.9). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 22 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | SciLitBench Llama 3.3 70B (Ollama build) | 78.6 |
| 2 | SciLitBench Llama 3.1 70B (Ollama build) | 75.5 |
| 3 | SciLitBench Qwen3 14B (Ollama build) | 74.9 |
| 4 | SciLitBench gpt-oss 120B (Ollama build) | 69.7 |
| 5 | SciLitBench Llama 3 70B (Ollama build) | 67.9 |
| 6 | SciLitBench Qwen3 32B (Ollama build) | 66.8 |
| 7 | SciLitBench Qwen3 8B (Ollama build) | 66.5 |
| 8 | SciLitBench gpt-oss 20B (Ollama build) | 62.8 |
| 9 | SciLitBench Llama 3.1 8B (Ollama build) | 60.9 |
| 10 | SciLitBench Mistral Large 123B (Ollama build) | 60.6 |
Interactive version: theaggregate.ai/benchmark?slug=scilitbench-title-and-abstract-screening · How It Works · Data refreshed daily, snapshot 2026-09-26.