Home / Companies / Weaviate / Blog / Post Details
Content Deep Dive

The Details Behind the Sphere Dataset in Weaviate

Blog post from Weaviate

Post Details
Company
Date Published
Author
Zain Hasan
Word Count
1,455
Company Posts That Month
5
Language
English
Hacker News Points
-
Post removed?
No
Summary

This article discusses the process of importing a large dataset into Weaviate using Apache Spark. The author provides detailed information on the hardware and software setup used for this task, including the use of Google Kubernetes Engine (GKE) nodes and the text2vec-huggingface vectorizer module in Weaviate. The article also covers performance metrics during the import process, such as batch duration and LSM store size. Additionally, it mentions future developments to improve memory usage at scale, including Vamana and HNSW+PQ technologies.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Kubernetes 1 1,156 140 65 -26%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.