Home / Companies / Onehouse / Blog / Post Details
Content Deep Dive

Diving into Uber's Cutting-Edge Data Infrastructure

Blog post from Onehouse

Post Details
Company
Date Published
Author
-
Word Count
1,841
Company Posts That Month
2
Language
English
Hacker News Points
-
Post removed?
No
Summary

At the Open Source Data Summit 2023, Girish Baliga, Director of Engineering at Uber, discussed the company's data infrastructure evolution, which enables Uber to process and analyze vast amounts of data efficiently across its global operations. Uber's data platform employs the Lambda architecture to manage both real-time and batch workloads, and it supports various analytics use cases including streaming, real-time, interactive, and batch analytics. Apache Hudi plays a crucial role in reducing data latency and improving data processing efficiency within this architecture. The platform uses Presto SQL as its primary query language, supported by Flink SQL and Spark SQL for more specialized needs, alongside programmatic APIs for complex scenarios. To enhance reliability and performance, Uber has implemented several technical customizations, such as smart query routing and multi-region deployments, and is transitioning towards a hybrid data environment with on-premises and cloud-based components. These innovations underscore Uber's commitment to leveraging open-source technologies to maintain a scalable and robust data infrastructure.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.