FinHardBench: leaderboard

Metric: Simulation pass rate (%; share of 99 single-shot generations per model, 33 financial FPGA module tasks x 3 trials, whose Verilog passes the self-checking testbench of its task in Icarus Verilog; the prompt holds only the natural-language specification; via OpenRouter at temperature 0.7). Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.661
2GPT-5.461
3Gemini 3.1 Flash Lite (Preview)46
4DeepSeek V3.225
5Mistral Large 323
6MiniMax-M2.719

Interactive version: theaggregate.ai/benchmark?slug=finhardbench · How It Works · Data refreshed daily, snapshot 2026-09-29.