Home / Companies / Onehouse / Blog / Post Details
Content Deep Dive

Scaling and Governing Robinhood’s Data Lakehouse

Blog post from Onehouse

Post Details
Company
Date Published
Author
-
Word Count
1,543
Company Posts That Month
5
Language
English
Hacker News Points
-
Post removed?
No
Summary

At the Open Source Data Summit 2023, Robinhood's data team, led by Balaji Varadarajan and Pritam Dey, detailed their implementation of a data lakehouse using Apache Hudi to handle exponential growth to a multi-petabyte scale. The system is built on a tiered architecture that efficiently manages over 10,000 data sources and supports various use cases from real-time streaming to analytics. This architecture facilitates robust data governance, ensuring GDPR compliance through mechanisms like efficient PII deletion, and maintains data freshness and access controls. Utilizing open-source technologies such as Debezium, Kafka, and Spark, the architecture allows for resource isolation, seamless scaling, and the separation of storage and compute, enabling Robinhood to stay competitive by efficiently managing data growth and compliance requirements.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.