PlantMarkerBench - Evidence Type (Maize): leaderboard
Metric: Evidence-type macro-F1 (%): macro-averaged F1 over the five evidence labels (expression, localization, function, indirect, noise) on the balanced 600-sentence Maize pilot split, times 100; open-weight models served by Ollama with the default prompt, OpenAI models with the direct prompt, temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around March 2027. 17 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 Mini | 40 |
Interactive version: theaggregate.ai/benchmark?slug=plantmarkerbench-evidence-type-maize · How It Works · Data refreshed daily, snapshot 2026-10-07.