mDenseOn with the mLateOn: Open Multilingual, Long-Context, and Code Retrieval Models
Blog post from Hugging Face
mDenseOn and mLateOn are two open-source multilingual retrieval models developed to enhance data retrieval across multiple languages and contexts, using a 2.8 billion-pair translate-train corpus, one of the largest to date. These models extend the successful English data recipe of DenseOn and LateOn to eight additional languages, focusing on overcoming the limitations of closed training data. mLateOn excels in multilingual tasks, outperforming mDenseOn by effectively generalizing to languages and scripts not seen during retrieval training, achieving high scores on MIRACL and MLDR benchmarks. The approach involves translating a curated English corpus into target languages to create multilingual and cross-lingual datasets, reinforcing the models' ability to transfer to languages outside the initial training set. Both models are publicly released with their datasets and training codes, underscoring the potential for open data recipes to compete with larger, closed datasets in multilingual retrieval tasks.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Model Fine-tuning | 33 | 896 | 206 | 76 | +18% |
| Vector Search | 5 | 2,031 | 414 | 136 | +6% |
| Data Pipeline | 1 | 519 | 185 | 75 | -1% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.