Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

Bekko Embedding: how small can a multilingual retrieval model be?

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Yuichi Tateno
Word Count
2,746
Company Posts That Month
73
Language
-
Hacker News Points
-
Post removed?
No
Summary

Bekko Embedding explores the development of compact multilingual retrieval models by significantly reducing the size of text embedding models while maintaining usable quality. The article describes two models, bekko-a8m and bekko-a25m, designed with active parameters of 7.67M and 24.93M, respectively, compared to larger models with billions of parameters. Despite their smaller size, these models perform competitively on the MMTEB Multilingual v2 benchmark, especially in retrieval tasks, and offer practical advantages such as faster inference times on modest hardware, including CPUs and even Raspberry Pi devices. Bekko's efficiency stems from pruning the mmBERT-small encoder, retaining key layers, and training on 1.1 billion multilingual pairs without using teacher models or distillation, all conducted on a single GPU. The models excel in scenarios where hardware resources are limited or browser-based deployment is required, providing a viable option for those seeking parameter-efficient, multilingual retrieval capabilities.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 44 2,031 414 136 +6%
LLM 2 7,115 1,261 236 +13%
AI Model Fine-tuning 1 896 206 76 +18%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.