Tool-Genesis (Direct) - Unit Tests (Soft): leaderboard

Metric: Unit-test score (0-1, times 100) over the standard tests: each test averages a key-path F1 and an embedding similarity between the generated tool's output and the expected output, single-pass generation (Direct), Tool-Genesis's tool-creation tasks over 86 executable MCP servers (508 tools, 9,441 unit tests): from an abstract requirement the model writes the tool interface schema and an executable MCP server implementation; temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around January 2028. 14 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.128.1#131
2GPT-4.126.1#240
3DeepSeek V3.219.5#198
4Qwen 3 32B17.6#424
5Qwen 3 14B15.4#524
6Kimi K214.4#236
7Qwen 3 235B A22B 2507 Instruct14.3#291
8GPT-4.1 Mini12.7#346
9Qwen 3 30B A3B 2507 Instruct9.3#464
10GPT-4o8.9#333
11Gemini 3 Flash (Preview)8.4#78
12Claude 3.5 Haiku0.7#553
13Qwen 3 8B0.1#667
14Qwen 3 4B0#823

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=tool-genesis-direct-unit-tests-soft · How It Works · Data refreshed daily, snapshot 2026-10-11.