Home / Companies / Onehouse / Blog / August 2024

August 2024 Summaries

5 posts from Onehouse

Filter
Month: Year:
Post Summaries Back to Blog
Many organizations utilize cloud data warehouses like Snowflake, Amazon Redshift, and Google BigQuery for data storage and analysis due to their tightly integrated systems, but this can lead to data lock-in and limited flexibility. Snowflake recently introduced support for Apache Iceberg to address these issues, offering options for managing data in an open table format while still using its query engine. However, these options come with limitations, such as restricted write access for other query engines and the need for Snowflake's catalog management. To overcome these challenges, Onehouse offers the Universal Data Lakehouse, which supports full interoperability across Apache Hudi, Apache Iceberg, and Delta Lake formats, allowing users to maintain data on their own cloud storage and access it with any query engine without extra costs. This solution enables seamless data sharing and management between platforms like Snowflake and Databricks, optimizing performance and compliance efforts without duplicating data.
Aug 29, 2024 1,389 words in the original blog post.
Generative AI is driving a rapid shift towards data lakehouse architecture, which efficiently supports AI applications by generating and managing vector embeddings. Onehouse has introduced a new AI vector embeddings generator that automates the creation of these embeddings, making AI projects more scalable for data practitioners. Vector embeddings represent unstructured data as mathematical constructs, facilitating their use in applications like semantic search and recommendation engines. Onehouse enables users to generate embeddings directly during data ingestion or transformation, utilizing models from providers such as OpenAI and Voyager AI. The generated embeddings are stored in Onehouse data lakehouse tables, leveraging the efficiencies of Apache Hudi for querying and updating vectors, thus ensuring data freshness without needing a separate freshness layer. This integration bridges the gap between vector databases, which are ideal for low-latency applications, and lakehouses, which offer scale and cost efficiency, creating a robust architecture for AI applications like large language models. Onehouse's open and interoperable framework allows seamless "reverse ETL" between data lakes and operational vector databases, supporting a variety of AI use cases, including Retrieval Augmented Generation and natural language processing.
Aug 22, 2024 881 words in the original blog post.
NOW Insurance is transforming the insurance sector for medical professionals by leveraging a lakehouse-powered data analysis and AI-driven innovations, facilitated by Onehouse's managed services. During an event hosted by Onehouse, Jonathan Sims, VP of Data and Analytics at NOW Insurance, and Andy Walner, Product Manager at Onehouse, discussed the adoption of a data lakehouse architecture powered by Apache Hudi, which significantly enhances NOW Insurance's data processing capabilities. This system allows for rapid, granular, and cost-effective data analysis, enabling the company to offer modern, personalized insurance products without the traditional bureaucratic hurdles. Key benefits for medical practitioners include a no-paperwork application process, fast insurance decisions, and customized coverage options. Onehouse's managed service offering has reduced engineering overhead, improved data freshness, and facilitated complex data analysis, ensuring that NOW Insurance can deliver real-time, data-driven products to its clients while maintaining compliance with industry standards. This partnership has enabled NOW Insurance to overcome engineering challenges related to frequent schema changes and inefficient data management, setting a new standard in the insurance industry.
Aug 16, 2024 1,573 words in the original blog post.
In the landscape of data architecture, the Universal Data Lakehouse (UDL) offers a flexible and open solution by decoupling storage from compute and centralizing data management, allowing users to choose query engines based on specific workloads rather than being restricted by architecture decisions. Onehouse implements this architecture, enabling seamless integration with various query engines and supporting a wide range of data processing needs from business intelligence to machine learning. The guide highlights the key considerations when selecting a query engine, such as manageability, scalability, cost, performance, and SQL support, while offering insights into popular engines like Amazon Athena, Google BigQuery, Snowflake, and open-source options such as ClickHouse and Trino. It emphasizes the importance of understanding unique workload requirements and suggests evaluating engines based on representative workloads to find the best fit. Ultimately, Onehouse's managed data lakehouse solution provides users with the flexibility to deploy multiple query engines and optimize data processes efficiently, eliminating traditional lock-in points and enhancing interoperability.
Aug 07, 2024 2,197 words in the original blog post.
Olameter, a leader in utility asset management, faced significant challenges in processing large volumes of XML data from electricity meters and transformers to predict outages, which initially took over six months for a year's worth of data due to inefficient custom .NET applications. To address this, they partnered with Onehouse, which developed a custom XML ingestion solution using Apache Hudi™ that enabled Olameter to process data incrementally and efficiently, reducing processing times from years to days. This collaboration also involved optimizing data structures for faster querying and analysis by flattening nested data and using geo-spatial clustering. Beyond data processing, Onehouse supports Olameter in building downstream pipelines and machine learning models, enhancing operational insights and service reliability. The partnership has allowed Olameter to achieve near real-time XML ingestion, significantly improve infrastructure management, and advance predictive maintenance capabilities, ultimately increasing quality of service and customer satisfaction in the utilities sector.
Aug 05, 2024 401 words in the original blog post.