June 2026 Summaries
2 posts from Mixedbread
Filter
Month:
Year:
Post Summaries
Back to Blog
Mixedbread Search describes how asymmetric quantization makes large-scale late-interaction retrieval more economical without substantially sacrificing ranking quality. Unlike single-vector embeddings, Wholembed v3 represents documents with hundreds of token-level vectors, improving precision but creating major storage and I/O costs across Silo’s index of more than 2.5 billion documents. The system addresses this by retaining query vectors in int8 precision while converting persistent document vectors into 1-bit sign representations, reducing raw multi-vector document storage from about 393 KiB to 12.28 KiB, a 32-fold reduction, while lowering average NDCG@10 only from 90.26 to 89.65. This approach exploits the fact that document vectors are stored, replicated, cached, and repeatedly retrieved, whereas queries are small and short-lived, making it more valuable to preserve query magnitude information than document precision. A specialized ARM scoring method evaluates int8 queries against packed binary document vectors without conventional multiplication for every dimension, contributing to a measured 3.82x speedup over fp32 scoring. Although fully binary query-and-document scoring is slightly faster, it causes a much larger quality loss, so int8-by-binary represents the preferred balance of storage efficiency, latency, and retrieval accuracy for production-scale multimodal late interaction.
Jun 29, 2026
1,805 words in the original blog post.
Neural retrieval models are argued to be limited less by the information they learn than by the scoring operators used to rank documents, with single-vector cosine similarity restricting what dense representations can express compared with late-interaction MaxSim approaches. The authors use sparse autoencoders (SAEs) to extract sparse “Latent Terms” from dense retriever activations without additional retrieval training, finding that these features form a roughly Zipfian distribution similar to natural-language vocabularies and include lexical, narrow semantic, and broad topical concepts. Because these sparse features resemble lexical terms, they can be indexed and ranked using BM25, producing retrieval performance that is competitive with or better than the originating single-vector models and, in some evaluations, comparable to SPLADE. On the LIMIT benchmark, Latent Terms substantially improve a dense model’s ability to retrieve documents requiring fine-grained attribute matching, supporting the view that dense retrievers encode relevance information not accessible through their usual scoring method. Experiments also indicate that this retrieval-ready sparse structure arises from retrieval-focused training rather than generic pretrained language-model representations, raising questions about better ways to extract and train such latent retrieval vocabularies.
Jun 02, 2026
3,601 words in the original blog post.