One Adapter, Both Modalities: Field Notes from Building and Serving a Multimodal Reranker
Blog post from Hugging Face
The development of LightOn-rerank, a 2B multimodal model designed to rerank both text passages and document pages, showcases significant advancements in reranking efficiency and effectiveness, particularly in the context of multimodal reranking, where both visual and textual data are considered. The model achieved a 62.66 NDCG@10 score on the ViDoRe V3 benchmark, surpassing previous models and demonstrating competitive performance on text-only benchmarks like BEIR. The key innovation lies in using a listwise approach, which contrasts with traditional pointwise methods by comparing multiple candidates simultaneously, allowing for richer cross-document comparisons. However, attempts to apply common text reranking speedup techniques, such as tournament scheduling or pointwise scoring, were ineffective due to the model's reliance on cross-document comparisons. The study emphasizes the importance of cross-document attention in enhancing reranking quality and suggests that scaling model size, while beneficial in listwise settings, does not yield similar gains in pointwise configurations. Additionally, the findings highlight that reducing the pool of candidates for reranking can significantly decrease computational costs with minimal impact on performance.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Model Fine-tuning | 10 | 402 | 99 | 46 | -46% |
| LLM | 6 | 3,751 | 612 | 168 | -39% |
| Vector Search | 6 | 1,111 | 224 | 91 | -41% |
| RAG | 1 | 619 | 146 | 64 | -38% |
| Reinforcement learning | 1 | 40 | 22 | 15 | -50% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.