jina-reranker-v3.5: Faster Listwise Reranking with Hybrid Attention and Self-Distillation
Blog post from Hugging Face
Jina AI has released jina-reranker-v3.5, a 0.6-billion-parameter listwise reranking model designed to improve retrieval quality and speed for enterprise search across general, multilingual, professional-domain, and semi-structured data. The model uses a hybrid attention architecture that combines sliding-window layers with strategically retained global-attention layers, enabling faster processing of long candidate lists while preserving cross-document context, and it is trained through a three-stage self-distillation process in which a full-attention model teaches an equal-sized sparse-attention version. On reported benchmarks, v3.5 achieved 63.20 nDCG@10 on BEIR, 74.11 on MIRACL, 70.95 on RTEB, and 48.3 on Struct-IR, with particularly large gains over its predecessor for field-constrained records, legal retrieval, and finance tasks. Its training data emphasizes hard retrieval cases in legal, medical, finance, multilingual, and structured settings, including synthetically generated examples that test numeric, date, equality, and logical constraints. Tests on an NVIDIA A100 found latency reductions of 1.22 times on short documents and 1.56 times on long documents compared with jina-reranker-v3, though the company notes remaining gaps against larger models in some legal, medical, low-resource-language, and structured-retrieval tasks, as well as inherent input-length and candidate-count limits of listwise reranking.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 7 | 525 | 92 | 52 | -74% |
| LLM | 2 | 1,189 | 251 | 109 | -83% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.