falcon-7B Instruct — benchmark results

Provider: TII. Released 2023-05-25. Access: Open.

Unified ELO 1149 ± 15, rank #1737 of 1776 rated models, from 106 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Trustworthy - Fairness100Trust Score (%)94
OpenEval - BoolQ79.69Exact Match (%)57.1
MMLU-by-task - TruthfulQA MC128.89Accuracy (%)52
MMLU-by-task - Machine Learning32.14Accuracy (%)48.7
MMLU-by-task - TruthfulQA MC244.07Accuracy (%)46.6
LLM Trustworthy - Stereotype87Trust Score (%)46
MMLU-by-task - Moral Scenarios25.14Accuracy (%)45.4
MMLU-by-task - Abstract Algebra29Accuracy (%)45.3
LLM Trustworthy - Ethics50.28Trust Score (%)44
MMLU-by-task - ARC Challenge42.15Accuracy (%)36.7
LLM Trustworthy - Adversarial43.98Trust Score (%)36
LLM Trustworthy - Toxicity39Trust Score (%)36

Interactive version: theaggregate.ai/model?slug=falcon-7b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.