Hermes-4-14B — benchmark results
Nous Research's open 14B Hermes 4 (September 2025), a Qwen3-14B-based hybrid reasoner with tool calling and deliberately neutral alignment. Provider: Nous Research. Released 2025-08-26. Access: Open.
Unified ELO 1527 ± 10, rank #695 of 1776 rated models, from 58 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LatamBoard - ASSIN2 RTE | 94.36 | Score (%) | 100 |
| LatamBoard - BLUEX | 74.27 | Score (%) | 100 |
| LatamBoard - ENEM Challenge | 83.07 | Score (%) | 100 |
| LatamBoard - FLORES Bidirectional | 47.41 | Score (%) | 100 |
| LatamBoard - FaQuAD NLI | 82.81 | Score (%) | 100 |
| LatamBoard - OAB Exams | 61.82 | Score (%) | 100 |
| LatamBoard - Portuguese Score | 90.81 | Score (%) | 100 |
| LatamBoard - Spanish PAWS | 66.35 | Score (%) | 100 |
| LatamBoard - Spanish Score | 67.28 | Score (%) | 100 |
| LatamBoard - Translation Score | 48.25 | Score (%) | 100 |
| LatamBoard - Spanish TeleIA | 76.19 | Score (%) | 98.5 |
| LatamBoard - Spanish XNLI | 49.44 | Score (%) | 93.9 |
Interactive version: theaggregate.ai/model?slug=hermes-4-14b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.