Home / Companies / Onehouse / Blog / June 2024

June 2024 Summaries

7 posts from Onehouse

Filter
Month: Year:
Post Summaries Back to Blog
Onehouse, a company focused on developing an open data architecture, recently raised $35 million in Series B funding led by Craft Ventures, with participation from existing investors Addition and Greylock Partners, to accelerate innovation and product development. The company has launched two new products: LakeView, a free lakehouse observability tool, and Table Optimizer, which automates data lakehouse optimizations. Onehouse emphasizes the importance of an open data architecture, advocating for decoupling storage and data management from the query engine to allow organizations to choose the right tools for various use cases, such as Generative AI or predictive ML, without data lock-in. The company is also working on table format and catalog interoperability to prevent data silos and compute lock-in, while promoting the use of open-source data processing frameworks like Apache Spark or Flink. Onehouse envisions providing a fully managed data lakehouse as a cloud service, enabling organizations to leverage both open-source solutions and managed services to meet their data needs without compromising on openness or flexibility.
Jun 26, 2024 1,597 words in the original blog post.
Onehouse introduces the Onehouse Table Optimizer, a product designed to enhance data lakehouse efficiency by automating table data layout optimizations and reducing the operational burden on engineering teams. By improving file sizing, clustering, compaction, and cleaning, the Table Optimizer can significantly boost query performance and decrease costs, achieving improvements of 2-10 times for customer tables. The tool integrates seamlessly with existing data ingestion pipelines, whether built with Onehouse or independently, and operates asynchronously to avoid bottlenecks, enabling efficient handling of large-scale workloads. As a managed service, Onehouse leverages its extensive experience to automate these optimizations, eliminating the need for complex infrastructure or manual tuning, and allowing organizations to focus on innovation while maintaining high performance and cost-effective operations.
Jun 26, 2024 953 words in the original blog post.
Onehouse has introduced LakeView, a free user interface for the Apache Hudi community that enhances data lakehouse observability with metrics, charts, and insights to optimize table performance without accessing base data files. As the first open-source data lakehouse technology, Apache Hudi has gained significant traction, and LakeView addresses the challenges in managing and optimizing data lakehouses by providing out-of-the-box observability, leveraging Onehouse's extensive experience with large data lakes. LakeView offers features such as interactive charts, metrics for monitoring table states, weekly email summaries, insights into data skew and file accumulation, and a searchable timeline for debugging issues. It enables teams to proactively monitor and optimize their data lakehouse environments by configuring alerts and tracking key metrics, making it easier to manage evolving data patterns and workloads.
Jun 26, 2024 814 words in the original blog post.
Onehouse provides organizations with a Universal Data Lakehouse model that enables flexible data management and analytics, including the ability to use Apache Hudi tables within Databricks. The blog outlines a detailed guide on setting up and configuring Apache Hudi in a Databricks environment, highlighting the ease of integrating Hudi with Databricks and the potential to mix and match table formats and query engines, thanks to tools like Apache XTable. The process involves creating compute instances, installing necessary libraries, and configuring the Databricks environment to support Hudi tables, allowing users to efficiently manage data operations. Additionally, the blog addresses how to potentially translate Hudi table metadata to Delta Lake format for use with Databricks’ Unity Catalog, although Unity Catalog currently supports only Delta Lake. This integration represents a streamlined approach to data management, offering developers the flexibility to handle data tasks with increased efficiency.
Jun 18, 2024 969 words in the original blog post.
OpenAI's ChatGPT 3.0 launch in 2022 significantly boosted the popularity of large language models and generative AI, leading to a focus on using data lakes for AI applications. Onehouse promotes a Universal Data Lakehouse vision, aiming to centralize enterprise data for enhanced AI utility. A critical aspect of generative AI applications is vector embeddings, which encode data into numerical vectors for similarity searches. Specialized vector databases like Pinecone and Milvus are expensive, especially when storing many vectors only some of which are necessary for specific use cases. Onehouse suggests a cost-efficient solution: generate and manage vector embeddings within a Universal Data Lakehouse, transferring only needed vectors to specialized databases for specific tasks. This approach balances cost and performance by using the lakehouse as the primary data repository and leveraging vector databases only when required. The integration of vector embeddings into data strategies through this method offers efficiency and scalability, crucial for modern AI-driven enterprises seeking to optimize resource allocation and maximize data utility.
Jun 13, 2024 1,960 words in the original blog post.
The blog post discusses the current landscape and future prospects of Apache Hudi within the context of the data lakehouse ecosystem, emphasizing its role as a robust open-source project with a strong community. It addresses the misconceptions about Hudi being merely a table format and highlights its innovative contributions, especially in incremental data processing and the open data lakehouse architecture. The article argues for the importance of open compute services to avoid vendor lock-in and showcases Hudi's achievements in supporting large-scale data operations across various cloud platforms. Additionally, it outlines the community's ongoing efforts to enhance Hudi's capabilities and interoperability with other table formats, while maintaining its unique strengths. The post encourages users to look beyond marketing narratives to make informed decisions based on technical merits and business needs, asserting that Hudi's open-source nature offers significant long-term advantages for data practitioners.
Jun 07, 2024 2,997 words in the original blog post.
An expert panel hosted by Onehouse discussed the evolution and challenges of data lakehouses, featuring insights from industry leaders at Uber, Walmart, and Robinhood. The discussion, moderated by Vinoth Chandar, highlighted the journey of adopting data lakehouses, the role of Apache Hudi, and the future of lakehouse technology. Panelists shared their experiences with large-scale implementations, including Uber's pioneering efforts in real-time transactional data and Walmart's migration to modern data architectures. They addressed technical challenges, regulatory compliance, and the benefits of open-source solutions like Apache Hudi. The conversation also explored the potential for increased interoperability and support for AI/ML use cases in future lakehouse systems, emphasizing the shift towards more relational database-like functionality and socio-technical data processing approaches.
Jun 04, 2024 1,993 words in the original blog post.