Home / Companies / Soda / Blog / January 2026

January 2026 Summaries

4 posts from Soda

Filter
Month: Year:
Post Summaries Back to Blog
Soda 4.0 introduces a significant evolution in data quality management by integrating AI, anomaly detection, and a new Data Contracts Engine within its unified platform, Soda Cloud. The new version replaces manual rule-setting processes with automated, AI-driven contract generation that allows data teams to define and enforce data quality standards through executable contracts, thereby streamlining collaboration between business and engineering teams. With support for various data sources such as Databricks, Snowflake, and BigQuery, Soda Core 4.0 shifts from individual checks to a contract-based syntax, simplifying data quality validation and maintenance. The platform's enhanced features include smarter anomaly detection, historical metrics analysis, and a Diagnostics Warehouse for in-depth data quality insights, aiming toward the vision of a self-driving data quality system.
Jan 28, 2026 1,534 words in the original blog post.
Soda Core is transitioning its license from the Apache License 2.0 to the Elastic License 2.0 (ELv2), maintaining its status as source-available and free to use, with new restrictions on commercial software that integrates Soda Core as a data quality engine. Users can continue employing Soda Core for internal purposes, including business use, without changes, provided it is not offered as a hosted or managed service to third parties. The ELv2 permits organizations to freely use, modify, and fork the source code, run it in production, and offer professional services related to it, such as consulting and support. However, the license prohibits offering Soda Core as a managed service to third parties, circumventing license key limitations, and removing licensing notices. The change is intended to support the development of source-available software while preventing third parties from offering Soda Core as a competing service. Further details and clarifications can be found in the full Elastic License 2.0 text and its FAQ.
Jan 27, 2026 290 words in the original blog post.
In 2026, organizations are prioritizing robust data foundations to support the reliable and responsible deployment of AI systems, shifting focus from AI experimentation to execution and emphasizing data quality management as a top priority. As AI models become more autonomous, the need for high-quality, well-governed data becomes critical, with organizations treating data quality metrics as leading indicators of AI return on investment. Automated observability and anomaly detection are increasingly integrated into data pipelines, enhancing data monitoring and enabling proactive problem resolution. Natural language processing-driven platforms are democratizing data quality management, allowing non-technical users to engage with data quality through intuitive interfaces. Data contracts and adaptive governance frameworks are facilitating clear expectations and accountability among teams, while data lineage and transparency become crucial for compliance and trust, especially amid growing regulatory pressures. Collectively, these trends reflect a balanced approach where leading organizations invest in both innovation and the fundamentals of data management to confidently scale AI while maintaining accountability and reducing risk.
Jan 15, 2026 1,674 words in the original blog post.
Data contracts are becoming increasingly important in distributed data ecosystems as a means to bridge the gap between data producers and consumers by establishing enforceable agreements that define how data should be structured, validated, and governed. These contracts aim to replace scattered, pipeline-specific checks with a unified approach that makes expectations explicit and reliable across teams and systems, similar to APIs in software development. By ensuring that key properties such as schema, nullability, and domain constraints are continuously verified against production data, data contracts transform informal guidelines into enforceable rules, thus creating a shared source of truth. They address systemic issues such as misaligned expectations, fragmented validation, and the lack of an authoritative source of truth by providing a formal interface that helps prevent unpredictable changes and downstream failures. Data contracts involve collaborative efforts between producers, who implement and enforce the contracts, and consumers, who define the requirements, while pipelines automatically validate each new batch of data before it reaches consumers. Tools like Soda support this contract-driven development by offering both code-based and no-code solutions to define, enforce, and monitor data quality expectations, facilitating collaboration between technical and non-technical users.
Jan 06, 2026 3,024 words in the original blog post.