FinInteract (Frontier Pilot): leaderboard
Metric: Accuracy (%; the stratified 50-instance pilot subset run for the cross-vendor panel; ReAct agent in the standard full-interaction mode (A+S+I): it may search the filing corpus, ask the simulated user (GPT-5, answering yes, no or I don't know from the hidden disambiguating context) yes-or-no questions, and answer; a final answer counts as correct only if it matches the intended interpretation, graded by GPT-4o-mini with a 1% numeric tolerance). Source: arxiv.org. Saturation forecast: Around 2029. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 5 | 44 |
| 2 | GLM-5.2 | 40 |
| 3 | Kimi K3 | 34 |
| 4 | Grok 4.5 | 32 |
| 5 | GPT-5.6 Sol | 30 |
| 6 | Gemini 3.5 Flash | 30 |
| 7 | DeepSeek V4 Flash | 24 |
Interactive version: theaggregate.ai/benchmark?slug=fininteract-frontier-pilot · How It Works · Data refreshed daily, snapshot 2026-09-26.