Home / Companies / Snowplow / Blog / April 2021

April 2021 Summaries

2 posts from Snowplow

Filter
Month: Year:
Post Summaries Back to Blog
In 2020, there was a notable increase in the adoption of cloud data warehouses, with Snowflake, Google BigQuery, and Amazon Redshift seeing significant growth as companies moved away from traditional on-premises systems to cloud-based solutions for better scalability and integration with other cloud services. This shift, part of a broader data warehouse modernization trend, allows organizations to efficiently scale their data operations and integrate data more flexibly, leading to enhanced data-driven decision-making and innovative product features such as personalized content and real-time recommendations. While traditional on-premises data warehouses require substantial physical infrastructure and management, cloud data warehouses offer on-demand scalability, seamless integration with tools like Snowplow, and the ability to handle growing data volumes from diverse sources. Organizations now face the challenge of selecting the right data warehouse by considering factors such as data types, pricing, and scalability to meet their specific needs. The concept of a "lakehouse" is emerging, which combines the features of data warehouses and data lakes, potentially simplifying data management by utilizing a single storage layer. Snowplow plays a crucial role in helping companies deliver and manage behavioral data across cloud data warehouses, facilitating real-time data availability and transformation to enhance data productivity.
Apr 28, 2021 1,860 words in the original blog post.
Data downtime, a term introduced by Monte Carlo to describe periods of inaccurate or incomplete data, poses significant challenges for companies reliant on data-driven decisions, as exemplified by a real incident at Acme where inaccurate data in a key report led to a loss of confidence in data among the leadership team. This issue highlights the importance of data observability, which offers transparency and control over data pipelines to quickly identify and resolve issues, thus minimizing data downtime. Unlike monitoring, which addresses known issues, data observability tackles unknown problems, providing comprehensive visibility to ensure data reliability. Snowplow's approach to data observability focuses on key metrics like throughput and latency to diagnose bottlenecks efficiently, aiming to create the most observable behavioral data pipeline for aligning technical data with business outcomes.
Apr 19, 2021 873 words in the original blog post.