GEITje-7B-ultra: benchmark results

Provider: Other. Access: Open.

Unified ELO 1358 ± 47, rank #2182 of 2656 rated models, from 19 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Grip on LLMs - Dutch Simplification - Amsterdam46.91SARI (0-100)100
Grip on LLMs - Dutch Simplification - INT Duidelijke Taal44.63SARI (0-100)100
Grip on LLMs - Dutch Simplification Average45.77SARI (0-100; mean of two datasets)100
Grip on LLMs - HonestCityBench - Narrow Expertise5Appropriate acknowledgement rate (%; LLM judge)91.7
Grip on LLMs - HonestCityBench - Outdated Information32Appropriate acknowledgement rate (%; LLM judge)70
Grip on LLMs - HonestCityBench - Insufficient Information35Appropriate acknowledgement rate (%; LLM judge)66.7
Grip on LLMs - HonestCityBench - Average26Appropriate acknowledgement rate (%; mean of five categories50
Grip on LLMs - HonestCityBench - Non-Text Modality18Appropriate acknowledgement rate (%; LLM judge)48.3
Open LLM Leaderboard - IFEval37.23Score36.1
Grip on LLMs - HonestCityBench - Incorrect Premise41Appropriate acknowledgement rate (%; LLM judge)30
Open LLM Leaderboard - BBH37.76Score21.8
Open LLM Leaderboard - MMLU-Pro20.11Score21.7

Interactive version: theaggregate.ai/model?slug=geitje-7b-ultra · How It Works · Data refreshed daily, snapshot 2026-09-19.