DualEvasion - Textual Evasion: leaderboard
Metric: Macro F1 (%; zero-shot classification of 505 earnings-call answers from 60 calls in 2023-2025 as direct or evasive from the transcript text, against labels from finance professionals; mean of the direct and evasive class F1). Source: arxiv.org. Saturation forecast: Around December 2026. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 87.4 |
| 2 | Gemini 3 Flash | 84.7 |
| 3 | Gemini 2.5 Flash | 75.1 |
| 4 | Llama 3.1 8B Instruct | 70 |
| 5 | Qwen 2 7B Instruct | 64.8 |
Interactive version: theaggregate.ai/benchmark?slug=dualevasion-textual-evasion · How It Works · Data refreshed daily, snapshot 2026-09-26.