Home / Companies / Mixedbread / Blog / Post Details
Content Deep Dive

Asymmetric Quantization: Near-Lossless Late Interaction Retrieval with 97% Storage Reduction

Blog post from Mixedbread

Post Details
Company
Date Published
Author
Aamir Shakir, Joel Dierkes, Rui Huang
Word Count
1,805
Company Posts That Month
2
Language
English
Hacker News Points
110
Post removed?
No
Summary

Mixedbread Search describes how asymmetric quantization makes large-scale late-interaction retrieval more economical without substantially sacrificing ranking quality. Unlike single-vector embeddings, Wholembed v3 represents documents with hundreds of token-level vectors, improving precision but creating major storage and I/O costs across Silo’s index of more than 2.5 billion documents. The system addresses this by retaining query vectors in int8 precision while converting persistent document vectors into 1-bit sign representations, reducing raw multi-vector document storage from about 393 KiB to 12.28 KiB, a 32-fold reduction, while lowering average NDCG@10 only from 90.26 to 89.65. This approach exploits the fact that document vectors are stored, replicated, cached, and repeatedly retrieved, whereas queries are small and short-lived, making it more valuable to preserve query magnitude information than document precision. A specialized ARM scoring method evaluates int8 queries against packed binary document vectors without conventional multiplication for every dimension, contributing to a measured 3.82x speedup over fp32 scoring. Although fully binary query-and-document scoring is slightly faster, it causes a much larger quality loss, so int8-by-binary represents the preferred balance of storage efficiency, latency, and retrieval accuracy for production-scale multimodal late interaction.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 2 1,918 398 137 -21%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.