Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

mDenseOn with the mLateOn: Open Multilingual, Long-Context, and Code Retrieval Models

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Raphael Sourty, Antoine Chaffin, Paulo Moura, and Amélie Chatelain
Word Count
5,088
Company Posts That Month
74
Language
-
Hacker News Points
-
Post removed?
No
Summary

mDenseOn and mLateOn are two open-source multilingual retrieval models developed to enhance data retrieval across multiple languages and contexts, using a 2.8 billion-pair translate-train corpus, one of the largest to date. These models extend the successful English data recipe of DenseOn and LateOn to eight additional languages, focusing on overcoming the limitations of closed training data. mLateOn excels in multilingual tasks, outperforming mDenseOn by effectively generalizing to languages and scripts not seen during retrieval training, achieving high scores on MIRACL and MLDR benchmarks. The approach involves translating a curated English corpus into target languages to create multilingual and cross-lingual datasets, reinforcing the models' ability to transfer to languages outside the initial training set. Both models are publicly released with their datasets and training codes, underscoring the potential for open data recipes to compete with larger, closed datasets in multilingual retrieval tasks.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Model Fine-tuning 33 975 221 80 +28%
Vector Search 5 2,241 449 143 +17%
Data Pipeline 1 530 192 77 +1%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.