Home / Companies / Onehouse / Blog / Post Details
Content Deep Dive

Notion's Journey Through Different Stages of Data Scale

Blog post from Onehouse

Post Details
Company
Date Published
Author
-
Word Count
1,877
Company Posts That Month
2
Language
English
Hacker News Points
-
Post removed?
No
Summary

In the Hudi Live Event, Notion's software engineers Thomas Chow and Nathan Louie detailed their evolution of data infrastructure in response to a 10x growth in data scale over three years, shifting from a single PostgreSQL database to a sharded setup, and ultimately adopting Apache Hudi for a universal data lakehouse architecture. This transformation was driven by the need to manage the rapid doubling of data every six months to a year, which challenged their previous systems and increased demands on data processing and analytics, especially with the introduction of generative AI features. The new architecture, integrating Postgres, Debezium CDC, Kafka, and Apache Spark, achieved significant cost savings of over a million dollars a year and improved performance, with historical Fivetran syncing times reduced from a week to two hours. This infrastructure supports the Q&A AI feature, allowing efficient processing of large data volumes and enabling real-time updates through a vector database, crucial for AI capabilities within Notion's platform.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.