Home / Companies / DataStax / Blog / Post Details
Content Deep Dive

Indexing All of Wikipedia on a Laptop

Blog post from DataStax

Post Details
Company
Date Published
Author
Jonathan Ellis
Word Count
1,581
Company Posts That Month
7
Language
English
Hacker News Points
-
Post removed?
No
Summary

Cohere has released a dataset containing all of Wikipedia chunked and embedded to vectors using their multilingual-v3 model. This makes creating a semantic, vector-based index of Wikipedia practical for an individual for the first time. The JVector library now supports indexing larger-than-memory datasets by performing construction-related searches with compressed vectors. By using Locally-Adaptive Quantization (LVQ) compression, it improves on previous methods and allows for faster searches while maintaining accuracy. This has enabled the indexing of all of English Wikipedia on a laptop, which was previously not practical.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 5 1,783 228 85 +36%
Real-time 1 2,587 688 208 +9%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.