Home / Companies / Soda / Blog / June 2025

June 2025 Summaries

10 posts from Soda

Filter
Month: Year:
Post Summaries Back to Blog
Appfire, a global software provider, has enhanced its data management capabilities by integrating Soda Core and Soda Cloud, which significantly reduced data scan times and improved real-time observability across its product portfolio. Facing challenges due to rapid growth and the complexity of managing data across over 100 diverse products, Appfire transitioned from a manual quality assurance process to a centralized, scalable data quality framework. The implementation of Soda's tools allowed Appfire to automate issue tracking, codify data quality standards, and establish a reliable governance system, improving data reliability and enabling better decision-making. As a result, Appfire has achieved a centralized view of data health, standardized governance policies, and a foundation for proactive quality management, ultimately setting a new standard for data reliability and trust within its digital product development.
Jun 30, 2025 1,190 words in the original blog post.
Soda's latest features, Metrics Observability and Collaborative Data Contracts, are designed to enhance data reliability within the Databricks ecosystem by allowing both technical and non-technical users to participate in data quality management without the need for coding. These features integrate seamlessly with Databricks, offering a no-code interface through Soda Cloud for business users and flexible integration paths for data engineers. Metrics Observability involves continuous monitoring of data health, using a proprietary metrics monitoring engine to detect anomalies and provide insights into data patterns, while Collaborative Data Contracts serve as agreements between data producers and consumers, defining data structure, quality expectations, and delivery guarantees. This system is designed to prevent data issues by shifting quality assurance earlier in the data lifecycle, facilitating collaboration, and ensuring data quality standards are met. By combining these features with Databricks' native tools, Soda provides a comprehensive approach to data quality, enabling organizations to detect and resolve data issues efficiently before they impact business operations.
Jun 19, 2025 3,247 words in the original blog post.
Soda's latest release, showcased at the Databricks Data & AI Summit, introduces a comprehensive integration with Databricks and other major data platforms, aiming to enhance data quality and observability. The update features the acquisition of NannyML to create a more intelligent data quality platform, the introduction of a Metric Monitors Dashboard for intuitive data analysis, and a proprietary anomaly detection algorithm that outperforms existing solutions. Additionally, Soda has launched Collaborative Data Contracts to bridge the gap between technical and non-technical users, offering both code and no-code editing modes. A new Soda Contract Language supports clear communication between data producers and consumers. The company emphasizes transparent pricing with a new model that eliminates hidden costs, and celebrates its community with the launch of The Soda Swag Store. These initiatives are designed to promote seamless collaboration, enhance data reliability, and provide a user-friendly experience for data teams and businesses alike.
Jun 16, 2025 1,877 words in the original blog post.
Soda has launched a new merch store to celebrate its recent platform upgrades, including AI-powered metrics observability, collaborative data contracts, and the acquisition of NannyML. The store offers items like a programmable macropad for data engineers, a book on metrics, and various branded merchandise, all aimed at engaging and appreciating the community that supports their open-source projects. Both Soda and NannyML emphasize their commitment to open-source development, with Soda Core and NannyML's libraries serving as foundational tools for data quality and model monitoring. The company is also preparing future releases, including Soda Core v4 and NannyML v1, which promise significant enhancements. The merch drop acts as a token of appreciation for the community's contributions and marks a milestone in Soda's journey, with an invitation to continue building and innovating together.
Jun 13, 2025 406 words in the original blog post.
Soda Collaborative Data Contracts offer a new approach for aligning data producers and consumers to prevent data issues before they reach production, emphasizing usability and flexibility. These contracts provide a unified interface for teams to define and enforce expectations, using either code-based YAML or a low-code UI, thus accommodating various workflows. They address common challenges of traditional data contracts by being adaptable and collaborative, bridging the gap between business and engineering teams. The tool allows anomalies to be turned into enforceable checks, promoting early detection and prevention of data quality issues, ultimately enhancing data integrity and operational efficiency. It integrates seamlessly with existing systems, offering fast deployment and evolution with data needs, and is available for use with the Soda CLI or UI, although self-serve account creation for Soda Cloud is temporarily paused pending major updates.
Jun 11, 2025 403 words in the original blog post.
Soda Data Observability is an AI-powered system designed for rapid anomaly detection across datasets, offering a quick setup and immediate results without extensive model training. It claims to be 70% more accurate than Prophet-based systems and can handle billions of data rows with minimal configuration. The tool addresses unexpected issues in data production, such as schema changes and volume shifts, with features like historical data backfill and comprehensive anomaly monitoring. Users can improve their data quality by transforming unidentified issues into tested expectations, while the system's explainable visualizations and feedback-aware detection enhance understanding and accuracy. Proven at scale, the system has already processed over a billion data rows swiftly on platforms like Databricks. Additionally, a promotional offer includes a chance to win a custom mechanical keyboard for those who sign up.
Jun 10, 2025 402 words in the original blog post.
Soda has introduced Soda Data Observability, an AI-powered system designed to detect anomalies across datasets quickly and efficiently, with a setup time of under five minutes and immediate results without extensive model training. This solution is reported to be 70% more accurate than Prophet-based systems and is capable of handling operations at a large scale, processing billions of rows swiftly. It offers features such as the ability to backfill up to a year of historical data, enabling immediate visibility into past anomalies, and provides explainable anomaly visualizations that detail expected ranges, impact, and trend history. The system is designed to improve data quality by helping users transition from unknown issues to tested expectations, enhancing pipeline resilience. Users are encouraged to schedule a demo or request a free account to explore further optimization of their data quality strategy, with an added incentive of a chance to win a custom mechanical keyboard for signing up within the week.
Jun 10, 2025 434 words in the original blog post.
Soda has acquired NannyML to create an advanced, context-aware data quality platform that aims to prevent issues before they develop into business problems and detect significant anomalies, allowing for comprehensive root cause analysis from data ingestion to automated decision-making. This partnership unites teams with a common goal of providing reliable systems that cater to modern AI and data infrastructure needs, addressing gaps in traditional data quality solutions that struggle with dynamic environments and complex real-time processes. NannyML's expertise in estimation-based performance monitoring and drift detection will enhance Soda's capabilities, offering smarter detection, context-aware alerting, and end-to-end observability, ensuring alignment between data and AI behaviors. The acquisition comes at a crucial time as data-driven systems become more complex and high-stakes, requiring tools that are impact-aware and AI-native. Soda's integration with NannyML will remain open-source and is set to expand, with new capabilities and product releases expected throughout the week to demonstrate the practical implementation of AI-first data quality solutions.
Jun 09, 2025 762 words in the original blog post.
Soda is a data quality platform designed to enhance real-time data observability and maintain reliable data pipelines, especially for organizations using scalable platforms like Databricks. It offers automated data quality checks through both a no-code UI and programmatic integration, allowing both technical and non-technical users to monitor and improve data reliability without needing to write code. The integration with Databricks is achieved through two main paths: using Databricks SQL Warehouse or PySpark with the soda-spark-df package, enabling data quality checks on Delta Lake tables or Spark DataFrames. These integrations facilitate early issue detection, real-time anomaly detection, and collaborative issue resolution, ensuring that data teams can effectively build trust in their data pipelines. The Soda platform supports scalability, automation, early detection of anomalies, and governance, ultimately providing a robust framework that integrates seamlessly with Databricks to reinforce data quality and observability.
Jun 05, 2025 1,513 words in the original blog post.
As data transitions from analytics to AI, the emphasis on data quality has become increasingly crucial, yet often overlooked compared to the hype around AI and big data. This shift has led to the adoption of preventive measures like "shifting left," which involves addressing data quality issues early in the data lifecycle to prevent downstream problems. Key strategies include implementing data contracts and testing, which formalize agreements between data producers and consumers, ensuring data meets specified standards and quality thresholds. These contracts draw from software engineering principles, fostering a collaborative culture by clearly defining data expectations, ownership, and automated enforcement mechanisms. They also enable testing at various stages of the data lifecycle, from development to runtime, to proactively manage data quality. The integration of observability tools and automated testing within data platforms supports this proactive approach, enhancing transparency and collaboration across data teams and ultimately embedding quality into the core of data operations.
Jun 03, 2025 3,967 words in the original blog post.