Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

mDenseOn with the mLateOn: Open Multilingual, Long-Context, and Code Retrieval Models

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Raphael Sourty, Antoine Chaffin, Paulo Moura, and Amélie Chatelain
Word Count
5,088
Company Posts That Month
73
Language
-
Hacker News Points
-
Post removed?
No
Summary

mDenseOn and mLateOn are two open-source multilingual retrieval models developed to enhance data retrieval across multiple languages and contexts, using a 2.8 billion-pair translate-train corpus, one of the largest to date. These models extend the successful English data recipe of DenseOn and LateOn to eight additional languages, focusing on overcoming the limitations of closed training data. mLateOn excels in multilingual tasks, outperforming mDenseOn by effectively generalizing to languages and scripts not seen during retrieval training, achieving high scores on MIRACL and MLDR benchmarks. The approach involves translating a curated English corpus into target languages to create multilingual and cross-lingual datasets, reinforcing the models' ability to transfer to languages outside the initial training set. Both models are publicly released with their datasets and training codes, underscoring the potential for open data recipes to compete with larger, closed datasets in multilingual retrieval tasks.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Model Fine-tuning 33 896 206 76 +18%
Vector Search 5 2,031 414 136 +6%
Data Pipeline 1 519 185 75 -1%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.