BTZSC — leaderboard
Zero-shot text classification benchmark for cross-encoders, embedding models, rerankers, and LLM classifiers.
Metric: Macro-F1 (%). Source: huggingface.co. Status: saturation imminent. 35 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Mistral Nemo Instruct (2407) | 66.97 |
| 2 | Qwen 3 8B | 66.49 |
| 3 | Qwen 3 4B | 64.86 |
| 4 | Phi-4 Mini Instruct | 43.09 |
| 5 | Llama 3.2 3B Instruct | 43.02 |
| 6 | Gemma 3 1B (IT) | 35.91 |
Interactive version: theaggregate.ai/benchmark?slug=btzsc · How the rankings work · Data refreshed daily, snapshot 2026-07-22.