FinMCP-Bench - Multi-Tool: leaderboard
Metric: Tool F1 (%): the harmonic mean of tool precision and tool recall between the set of MCP tools the model calls and the reference set (order and grouping ignored), with the model acting as the agent of a real fund-advisory app over its 65 financial MCP tools and each turn answered after the gold conversation history, on the 249 multi-tool samples (one turn needing several sequential or parallel tool calls; real log samples and chain-synthesized ones, expert-reviewed); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 6 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Qwen 3 235B A22B 2507 (Thinking) | 69.42 | #253 (Qwen 3 235B A22B 2507) |
| 2 | Qwen 3 30B A3B 2507 (Thinking) | 60.73 | #366 (Qwen 3 30B A3B 2507) |
| 3 | DeepSeek R1 | 52.36 | #245 |
| 4 | Qwen 3 4B 2507 (Thinking) | 50.23 | #525 (Qwen 3 4B 2507) |
| 5 | GPT-OSS-20B | 38.54 | #499 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=finmcp-bench-multi-tool · How It Works · Data refreshed daily, snapshot 2026-10-11.