Home / Companies / Onehouse / Blog / September 2023

September 2023 Summaries

3 posts from Onehouse

Filter
Month: Year:
Post Summaries Back to Blog
Onehouse and CelerData have partnered to offer a streamlined data integration experience that bridges the gap between raw data and business value through an integrated data infrastructure that leverages Apache Hudi and StarRocks. This collaboration provides users with a faster and more unified data ingestion and analytics process, allowing automatic scaling, security, maintenance, and optimization of data using just a web browser and email address. The integration supports data ingestion from various sources, such as databases and event streams, and consolidates data into a centralized repository, enabling low-latency SQL analytics and saving time and costs with automated data management tasks. Users can explore these benefits through a 30-day free trial of the Onehouse and CelerData Cloud, facilitating the deployment of a data lakehouse and SQL querying with ease.
Sep 14, 2023 386 words in the original blog post.
In the second part of his talk at Data Council Austin in March 2022, Onehouse CEO Vinoth Chandar delves into the comparison of data warehouses and data lakehouses, focusing on their capabilities and price/performance attributes. He notes that while both architectures support core functionalities like updates and transactions, data warehouses often offer more complete relational database features, whereas lakehouses cater to higher-scale data needs and require hands-on management. Chandar highlights that the choice between warehouses and lakehouses hinges on workload requirements, with warehouses offering managed services and lakehouses providing flexibility but necessitating self-management. He discusses the evolving landscape of query optimization, the importance of benchmarking workloads to determine cost-effectiveness, and the advantages of a lakehouse-first architecture, which supports openness, interoperability, and a unified data source for both BI and data science teams. The talk concludes with best practices for leveraging lakehouses, emphasizing their ability to offer cost-effective ETL solutions while maintaining high performance and adaptability to emerging technologies like ML and streaming.
Sep 12, 2023 2,827 words in the original blog post.
Vinoth Chandar's talk at Data Council Austin 2022 explores the evolution and future of data infrastructure, focusing on data warehouses, data lakes, and the emerging data lakehouse architecture, which he helped develop at Uber. As the chair of the Apache Hudi project, Chandar emphasizes the significance of a lakehouse-first approach that integrates diverse engines for various analytics workloads. He outlines how lakehouses can address the limitations of data lakes by incorporating transactional capabilities, thus offering a balanced and future-proof data infrastructure solution. Comparing the architectural designs, capabilities, and cost-performance aspects of data warehouses and lakehouses, Chandar advocates for openness in data formats to promote interoperability and innovation across the ecosystem. He highlights the maturation of cloud warehousing and the advantages of separating storage and compute, while cautioning against over-reliance on vendor-specific solutions, emphasizing the potential of open-source projects like Hudi to maintain flexibility and community-driven development.
Sep 06, 2023 2,647 words in the original blog post.