Home / Companies / Snowplow / Blog / January 2016

January 2016 Summaries

1 posts from Snowplow

Filter
Month: Year:
Post Summaries Back to Blog
Bad data, often overlooked in digital analytics, poses significant challenges by undermining data reliability and the insights derived from it. Despite its unglamorous nature, addressing bad data is crucial, as it can erode organizational confidence in data sources, making it difficult to derive and socialize insights. The Snowplow pipeline offers a structured approach to managing bad data by making data quality auditable at every stage, ensuring that no data is dropped, and allowing for reprocessing of erroneous data. This system identifies the root causes of data issues—missing and inaccurate data—while employing methods such as self-describing data schemas and intelligent use of queues and autoscaling to minimize data loss. By retaining and analyzing bad data in a separate "bad rows" repository, Snowplow enables ongoing monitoring and resolution of data quality issues, contrasting with the black-box nature of many commercial analytics solutions that often obscure such problems.
Jan 07, 2016 2,986 words in the original blog post.