Notion's Journey Through Different Stages of Data Scale
Blog post from Onehouse
In the Hudi Live Event, Notion's software engineers Thomas Chow and Nathan Louie detailed their evolution of data infrastructure in response to a 10x growth in data scale over three years, shifting from a single PostgreSQL database to a sharded setup, and ultimately adopting Apache Hudi for a universal data lakehouse architecture. This transformation was driven by the need to manage the rapid doubling of data every six months to a year, which challenged their previous systems and increased demands on data processing and analytics, especially with the introduction of generative AI features. The new architecture, integrating Postgres, Debezium CDC, Kafka, and Apache Spark, achieved significant cost savings of over a million dollars a year and improved performance, with historical Fivetran syncing times reduced from a week to two hours. This infrastructure supports the Q&A AI feature, allowing efficient processing of large data volumes and enabling real-time updates through a vector database, crucial for AI capabilities within Notion's platform.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.