February 2021 Summaries
2 posts from Soda
Filter
Month:
Year:
Post Summaries
Back to Blog
Soda SQL is an open-source tool designed to enhance data quality by focusing on metric collection, data testing, and data monitoring for SQL-accessible data. Developed as part of Soda's strategy to provide robust data management tools, it empowers data engineers to implement test-driven development principles in their workflow, addressing the need for reliable data pipelines and protection against silent data issues. Soda SQL operates using industry-standard YAML configuration files, allowing engineers to define and audit data tests and metrics efficiently. The tool integrates seamlessly into existing data environments by using SQL, eliminating the need to transfer data for testing. Complementing Soda SQL is the Soda Data Monitoring Platform, which offers real-time insights through a no-code interface, enabling data teams to collaborate and manage data quality effectively. Soda's commitment to open-source development aims to foster community engagement and expand support to streaming and dataframes, ultimately advancing data observability across technology stacks.
Feb 12, 2021
1,112 words in the original blog post.
Dremio's Subsurface LIVE Winter Edition, a cloud data lake conference, showcased the rapidly evolving landscape of data management, emphasizing the increasing importance of data monitoring and the pivotal role of data engineers. The event highlighted the shift from traditional client-server architectures to cloud-based, open-source solutions to handle large datasets, as noted by Tomer Shiran in his keynote. This transition underlines the urgent need for on-demand data availability and the challenges of maintaining data quality across the data pipeline. The conference also marked the introduction of Soda SQL, an open-source project by Soda, emphasizing Test-Driven Development (TDD) principles to ensure data quality. Attendees collectively acknowledged the necessity of monitoring, testing, and validating data before it reaches the user, reflecting the growing pressure on data engineering teams to deliver consistent, analytics-ready data from a multitude of sources. Despite the challenges, the event fostered a sense of community and collaboration among data professionals, with discussions on integrating software engineering practices into data workflows.
Feb 05, 2021
762 words in the original blog post.