Home / Companies / Qdrant / Blog / October 2026

October 2026 Summaries

3 posts from Qdrant

Filter
Month: Year:
Post Summaries Back to Blog
No summary generated yet.
Oct 08, 2026 3,890 words in the original blog post.
Jev, a TypeSafe decision model served through OpenRouter, is presented as a generalized classifier that accepts runtime labels, produces constrained outputs, and aims to provide a faster, less costly alternative to generative LLM prompting for classification tasks. Experiments across search and retrieval found that Jev improved reranking performance on NFCorpus, with its single-request scoring method offering nearly the benefit of a ten-request iterative approach and outperforming tested local cross-encoders. In e-commerce search, Jev reranking increased relevance but could make results more repetitive, while combining its relevance scores with Maximal Marginal Relevance reduced near-duplicates and retained relevance above the hybrid-search baseline. For query understanding, Jev-based product-category routing improved Amazon-C4 results when high-confidence predictions filtered candidates and moderate-confidence predictions boosted them, although ambiguous short queries such as “apple” showed that confident classification can still harm results. Jev was also used for semantic document chunking by identifying topic changes between sentences, producing improvements over fixed, recursive, and embedding-based splitters on QASPER evidence retrieval, though at API cost. The authors recommend semantic chunking as an accessible RAG application, Jev-plus-MMR for product-search diversity, and cautious testing of taxonomy filters, concluding that generalized decision models may be valuable where their quality, latency, and cost trade-offs are acceptable.
Oct 08, 2026 2,197 words in the original blog post.
Google DeepMind’s EmbeddingGemma 2 is an open multimodal embedding model that maps text, code, images, video, and audio into a shared 768-dimensional space, with a 270-million-parameter text pathway, an 8K-token context window, and support for Matryoshka Representation Learning, which allows vectors to be shortened without re-embedding data. Qdrant tested memory-saving approaches for its vectors across five text-retrieval datasets, comparing reduced dimensionality and 1-bit TurboQuant compression against exact 768-dimensional float32 search. Full-size 768-dimensional vectors quantized to 1 bit reduced vector RAM by roughly 30 times while retaining 99.0% of baseline retrieval quality without rescoring and 99.7% with rescoring, while a 256-dimensional 1-bit configuration with fourfold oversampling and rescoring used 77 times less vector RAM while retaining 94.5% quality. The tests suggest that retaining more dimensions while using stronger quantization generally performs better than reducing vector length first, and recommend beginning with full-size 1-bit vectors without rescoring before evaluating rescoring or dimensional reduction based on an application’s relevance, latency, and memory requirements.
Oct 06, 2026 664 words in the original blog post.