Diving into Uber's Cutting-Edge Data Infrastructure
Blog post from Onehouse
At the Open Source Data Summit 2023, Girish Baliga, Director of Engineering at Uber, discussed the company's data infrastructure evolution, which enables Uber to process and analyze vast amounts of data efficiently across its global operations. Uber's data platform employs the Lambda architecture to manage both real-time and batch workloads, and it supports various analytics use cases including streaming, real-time, interactive, and batch analytics. Apache Hudi plays a crucial role in reducing data latency and improving data processing efficiency within this architecture. The platform uses Presto SQL as its primary query language, supported by Flink SQL and Spark SQL for more specialized needs, alongside programmatic APIs for complex scenarios. To enhance reliability and performance, Uber has implemented several technical customizations, such as smart query routing and multi-region deployments, and is transitioning towards a hybrid data environment with on-premises and cloud-based components. These innovations underscore Uber's commitment to leveraging open-source technologies to maintain a scalable and robust data infrastructure.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.