Asymmetric Quantization: Near-Lossless Late Interaction Retrieval with 97% Storage Reduction
Blog post from Mixedbread
Mixedbread Search describes how asymmetric quantization makes large-scale late-interaction retrieval more economical without substantially sacrificing ranking quality. Unlike single-vector embeddings, Wholembed v3 represents documents with hundreds of token-level vectors, improving precision but creating major storage and I/O costs across Silo’s index of more than 2.5 billion documents. The system addresses this by retaining query vectors in int8 precision while converting persistent document vectors into 1-bit sign representations, reducing raw multi-vector document storage from about 393 KiB to 12.28 KiB, a 32-fold reduction, while lowering average NDCG@10 only from 90.26 to 89.65. This approach exploits the fact that document vectors are stored, replicated, cached, and repeatedly retrieved, whereas queries are small and short-lived, making it more valuable to preserve query magnitude information than document precision. A specialized ARM scoring method evaluates int8 queries against packed binary document vectors without conventional multiplication for every dimension, contributing to a measured 3.82x speedup over fp32 scoring. Although fully binary query-and-document scoring is slightly faster, it causes a much larger quality loss, so int8-by-binary represents the preferred balance of storage efficiency, latency, and retrieval accuracy for production-scale multimodal late interaction.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 2 | 1,918 | 398 | 137 | -21% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.