Bekko Embedding: how small can a multilingual retrieval model be?
Blog post from Hugging Face
Bekko Embedding explores the development of compact multilingual retrieval models by significantly reducing the size of text embedding models while maintaining usable quality. The article describes two models, bekko-a8m and bekko-a25m, designed with active parameters of 7.67M and 24.93M, respectively, compared to larger models with billions of parameters. Despite their smaller size, these models perform competitively on the MMTEB Multilingual v2 benchmark, especially in retrieval tasks, and offer practical advantages such as faster inference times on modest hardware, including CPUs and even Raspberry Pi devices. Bekko's efficiency stems from pruning the mmBERT-small encoder, retaining key layers, and training on 1.1 billion multilingual pairs without using teacher models or distillation, all conducted on a single GPU. The models excel in scenarios where hardware resources are limited or browser-based deployment is required, providing a viable option for those seeking parameter-efficient, multilingual retrieval capabilities.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 44 | 2,031 | 414 | 136 | +6% |
| LLM | 2 | 7,115 | 1,261 | 236 | +13% |
| AI Model Fine-tuning | 1 | 896 | 206 | 76 | +18% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.