May 2023 Summaries
52 posts from Metaplane
Filter
Month:
Year:
Post Summaries
Back to Blog
Data security refers to protecting data from unauthorized access, use, theft or destruction. It encompasses a range of technologies, processes and policies designed to safeguard data from potential threats. Data privacy specifically deals with protecting personal information from intrusive or harmful use. Ensuring data security is important because insecure data can lead to damaged reputation, loss of revenue or customers, regulatory fines, and litigation. Examples of insecure data include exposure of sensitive data, manipulation of data, and loss of data. Metrics used to measure data security include data encryption coverage, user authentication attempts, and time to breach detection. To ensure data security, best practices such as access control and regular backups should be implemented.
May 29, 2023
837 words in the original blog post.
Data usability refers to the ease with which data can be accessed, understood, and used to inform business decisions. It is one of ten dimensions of data quality, along with elements like accuracy, validity, completeness, and consistency. Poor data usability can lead to inconsistent data formats, inadequate documentation, and poor-quality data, all of which can hinder decision-making processes.
To measure data usability, one can use table documentation, monitor query successes, and conduct user surveys. To ensure data usability, it is necessary to actively maintain data quality using real-world metrics such as standardizing data formats, implementing an anomaly detection system, and providing proper documentation. Data observability tools can also help enhance data usability initiatives by providing anomaly detection for key objects and related queries.
May 29, 2023
782 words in the original blog post.
Data Validity is a crucial aspect of data quality that ensures accuracy and relevance of data used for decision-making in businesses. It refers to the degree to which business rules or definitions are accurately represented. Invalid data can be caused by various issues such as data entry errors, system glitches, or intentional falsification. To measure data validity, metrics like completeness rate, accuracy rate, and timeliness rate are commonly used. Ensuring data validity involves using data validation rules, utilizing anomaly detection tools, and implementing continuous monitoring and validation of data quality.
May 29, 2023
613 words in the original blog post.
Data relevance is the degree to which data provides insight into a real-world problem or purpose being addressed and contributes to the overall understanding of the business. It is one of the ten dimensions of data quality, which also includes completeness, consistency, accuracy, and timeliness. Irrelevant data can lead to poor decision-making and damage a company's reputation. To measure data relevance, teams can use metrics such as data usage, time to analysis, and feedback from users. Ensuring data relevance involves identifying user needs, monitoring trends over time, and using data observability tools to assess the usefulness of data.
May 29, 2023
742 words in the original blog post.
Data reliability is an essential aspect of data quality that ensures trust in data for operational and decision-making purposes. Unreliable data can lead to costly mistakes, while reliable data provides a competitive edge. Common examples of unreliable data include incomplete, inconsistent, or missing data. To measure data reliability, metrics such as completeness rate, consistency rate, and accuracy rate can be used. Ensuring data reliability involves establishing data quality policies, utilizing anomaly detection, and choosing accessible data tools. By understanding and maintaining data reliability, businesses can make better decisions based on accurate and consistent data.
May 29, 2023
701 words in the original blog post.
Data completeness is an important aspect of data quality, which refers to the absence of missing information in a dataset. It has significant implications for business operations and decision-making processes. Incomplete data can lead to missed opportunities or incorrect conclusions that could negatively impact the organization. Ensuring data completeness involves measuring it against a complete mapping, tracking null values, satisfying constraints, and validating input mechanisms. Anomaly detection is one method to identify missing data in real-time, helping organizations maintain high levels of data quality.
May 28, 2023
673 words in the original blog post.
Data freshness is one of ten dimensions of data quality, and refers to the timeliness of data in relation to the present moment. It is crucial for businesses that rely on operational or decision-making purposes, as stale data can lead to costly mistakes or missed opportunities. To measure data freshness, various methods such as comparing latest timestamps against the present moment, verifying differences between a destination and source system, or corroborating with other pieces of data can be used. Anomaly detection is one way to ensure data freshness by identifying unexpected values or events in a dataset.
May 28, 2023
691 words in the original blog post.
Data accuracy is crucial for businesses as it impacts their bottom line and can lead to negative consequences if inaccurate. It is one of the ten dimensions of data quality, which refers to how well data describes the real world. Poor data accuracy can result from issues such as accidental manual data entry or mismatched geographies. To measure data accuracy, businesses can compare their data against a reference set, corroborate it with other data, and use anomaly detection software to identify unexpected values or events in a dataset. Ensuring data accuracy helps prevent negative downstream impacts on various uses of data, such as artificial intelligence and data analytics.
May 28, 2023
778 words in the original blog post.
Data consistency is one of ten dimensions of data quality, ensuring that two or more values in different locations are identical. It has a significant impact on business success and decision-making processes. Inconsistent data can lead to low data integrity and incorrect decisions based on misleading information. To measure data consistency, you can track the number of passed checks for uniqueness of values, entities, corroboration within the system, or referential integrity maintenance. Anomaly detection is a method to ensure data consistency by identifying unexpected values or events in a dataset. Proper data management and tracking are crucial to maintaining data consistency and ensuring accurate business decisions.
May 26, 2023
1,004 words in the original blog post.
As data becomes increasingly vital in business strategy, organizations must effectively manage, process, and store their data. Two popular architectures for this purpose are Data Mesh and Data Fabric. While Data Mesh emphasizes a decentralized approach to data management with each domain managing its own infrastructure, Data Fabric promotes seamless integration of diverse data sources into a single layer. The choice between these two approaches depends on factors such as the complexity of data models, data accessibility requirements, existing data infrastructure, and the organization's preference for decentralized or unified data management. Regardless of the chosen architecture, data observability is crucial in ensuring high-quality and consistent data across the entire platform.
May 25, 2023
1,660 words in the original blog post.
Data modernization is the process of updating an organization's data infrastructure to keep up with new technology and changing business needs. It involves improving data quality, accuracy, flexibility, real-time insights, efficiency, cost savings, and customer experience. Modernizing a data stack can help businesses make confident, timely, data-backed decisions to drive the business forward. The process includes identifying pain points in the current data infrastructure, evaluating modern data tools, making a data modernization plan, implementing tools, and automating data pipelines. Challenges of data modernization include resistance to change, skills gap, security and compliance concerns, integration issues, and high up-front costs. Trends in data modernization include adoption of cloud data platforms, focus on data protection, privacy, and governance, automated data processing with AI and ML, and real-time data processing.
May 25, 2023
2,406 words in the original blog post.
Data monitoring and data observability are two techniques used to manage complex data stacks and ensure their reliability and accuracy. While both serve the same goal, they differ in when and how they're used. Data monitoring primarily helps identify potential issues or disturbances in the data stack, while data observability provides tools for gaining insight into what's happening with any data issues and preventing future incidents. Data observability plays a crucial role in the modern data stack by analyzing constantly moving data to provide feedback about daily ad hoc decisions, ensuring that the data reflects business realities.
May 25, 2023
881 words in the original blog post.
Poor data quality costs businesses an average of $13 million per year, with the impact varying across industries. Ensuring data quality is crucial for organizations and requires understanding its role in specific business contexts, setting up data management and governance practices, and using appropriate tools to prevent and troubleshoot issues. The cost of low-quality data depends on factors such as the industry and how it's used within a company. Businesses typically use data in four ways: not at all, for operations, to inform strategy, or as a product. A three-part framework can help identify how data quality impacts business performance. To prevent and troubleshoot data quality issues, organizations should focus on the root cause of problems and utilize data observability tools to monitor dimensions of data quality and shorten time-to-detection and time-to-resolution.
May 25, 2023
1,802 words in the original blog post.
Poor data quality can lead to lost revenue, reduced operational efficiency, and bad business decisions. To maintain high-quality data, businesses should identify the most important dimensions, track metrics against those dimensions, and establish SLAs on those metrics. Two guidelines for meaningful data metrics are measuring what matters and making metrics actionable. Practices that can improve data quality include preventing issues that can be prevented, investing in data infrastructure tooling, checking metrics deeply and often, aligning data quality with stakeholder impact, and arming the data team with data quality tools.
May 24, 2023
2,462 words in the original blog post.
Data observability is the practice of monitoring and troubleshooting data pipelines and infrastructure to ensure accurate, complete, and consistent data. Machine learning observability focuses on monitoring and understanding machine learning models' behavior and performance. Both require continuous monitoring and utilize metrics monitoring and anomaly detection techniques.
May 24, 2023
931 words in the original blog post.
Both Data Mesh and Data Lake architectures have their advantages and drawbacks, and selecting the right one for your organization should be a thoroughly considered choice. A Data Mesh is a novel data platform design paradigm that emphasizes domain-driven decentralized data management, self-serve data infrastructure, and a federated governance model. On the other hand, a Data Lake is a centralized repository for storing raw data that can be processed and analyzed for different business purposes. The primary difference between these two architectures lies in their design principles: Data Mesh emphasizes a decentralized approach to data management, while Data Lake takes a centralized approach. Factors such as the complexity of data models, existing data infrastructure, transformation needs, and data observability must be taken into consideration when deciding between the two approaches.
May 24, 2023
1,888 words in the original blog post.
Data catalogs are organized and searchable inventories of an organization's data assets that provide a comprehensive view of available data, its lineage, and metadata. They help users find the right data at the right time by offering a single point of truth for data, making it easy to search, discover, and access various data assets in a timely and efficient manner. When evaluating data catalogs, consider factors such as data coverage, search and discovery capabilities, data lineage and metadata management, collaboration and community aspects, and scalability and integration capabilities. Some top data catalogs include Alation, Atlan, and Castor.
May 24, 2023
965 words in the original blog post.
Data observability is the degree of visibility into your data at any point in time. It involves collecting metadata about the properties and relationships between data, monitoring changes, and presenting actionable insights. Key aspects include the number of closed-ended questions a data team can answer, insight into the data itself rather than just the data system, and historical baselines for comparison. Data observability differs from data quality, data monitoring, and data testing in various ways, such as its focus on solving multiple problems and providing coverage across the entire data stack. It is important because it helps data teams provide high-quality data that businesses need to make better decisions and take effective actions.
May 24, 2023
2,136 words in the original blog post.
Data lineage is the record of the movement of data from its origin to its end destination, providing valuable information about where the data came from, how it was transformed, and where it is being stored. Both forward and backward lineage offer value in understanding the flow of data. Automated data lineage improves data governance, speeds up issue resolution, enhances data quality, promotes better collaboration and communication among teams, and supports machine learning applications. Proper data observability relies on data lineage, ensuring transparency into data flow and reducing the risk of errors.
May 24, 2023
1,133 words in the original blog post.
Data products refer to the use of data for decision making or problem solving within a business context. They have evolved from traditional analytics, such as reports and dashboards, to include advanced technologies like machine learning and AI APIs. The scope of data products is expanding rapidly, with companies recognizing their potential in driving growth and innovation. Data product management is emerging as a career field, handling the lifecycle of these products using traditional product thinking. Treating data products as actual products offers several benefits, including market research, user acceptance testing, go-to-market strategy, and iterative updates to ensure continued usefulness to the business.
May 24, 2023
1,054 words in the original blog post.
Microsoft Fabric is a new offering in data storage and management from Microsoft, integrating features such as a data lake, multiple compute setups, advanced data governance, and a wide application surface area. Key components include OneLake, which serves as the central data lake; versatile compute options including T-SQL, Spark, KQL, and Analysis Services; a unified security model for data management and governance; broad application scope across workloads such as data engineering, analysis, and science; and a flexible pricing model.
May 24, 2023
573 words in the original blog post.
Data Mesh is an approach to data management that shifts from centralized data platforms to domain-oriented, decentralized data management. It breaks down data bottlenecks and silos within an organization, allowing each domain team to take full ownership of their domain data. This results in increased scalability, faster insights, and a more optimized data-driven decision-making process. Key concepts include domain-oriented decentralized data management, self-serve data platforms and APIs, and discoverability and accessibility of datasets. Implementing Data Mesh effectively requires adherence to best practices such as establishing clear interfaces and standards for data sources and pipelines, enforcing data governance and domain ownership, ensuring data quality and metadata, implementing access controls and automated processes, promoting a culture of collaboration and communication, investing in proper training and skill development, ensuring data discoverability, planning for scalability, iterating and evolving. Some real-life examples of companies successfully implementing Data Mesh include ING, Zalando, and Intuit.
May 23, 2023
1,642 words in the original blog post.
The future of data analytics and data science is set to be shaped by automation, machine learning, LLMs, data quality, and democratization of data. Automation through LLMs can streamline data workflows, enabling faster insights and improved efficiency. Data observability tools are crucial for maintaining high-quality data, while democratizing data fosters a data-driven culture that encourages informed decision-making and innovation. These elements will be key in navigating the future of data analytics and science.
May 23, 2023
777 words in the original blog post.
Data Observability is an emerging concept borrowed from control theory that measures how well internal states of a system can be inferred from its external outputs. It has gained popularity in software systems and is now spreading to data systems. Software observability tools, such as Datadog, AppDynamics, New Relic, Grafana, Splunk, and Sumo Logic, have transformed the world of software by providing a centralized view across systems, enabling easier debugging, and improving overall system performance.
Data Observability is similar to Software Observability in that issues compound over time, are disruptive, require historical data for identification, and strive for reliable systems that prevent issues from occurring at the source and self-heal when issues do occur. However, Data Observability differs in its focus on data, which has weight, structure, and history, unlike software systems that have minimal marginal cost of replication, are interchangeable, and increasingly ephemeral.
Data observability tools monitor machine to machine interactions as well as many machine to person interactions, making it more complex to understand the impact of issues and communicate them effectively. The ecosystem of data observability platforms is still in its early stages, with various players approaching the problem from different perspectives.
May 23, 2023
2,652 words in the original blog post.
Data observability plays a crucial role in ensuring compliance with data governance policies as the data industry becomes more regulated and monitored. It refers to understanding what is happening within a software system or application based on its outputs, allowing data teams to monitor, troubleshoot, and ensure data quality across their stack. Implementing data observability tools adds transparency, reliability, and integrity, enabling data teams to adapt and evolve sustainably. Data governance policy is critical for modern data-driven organizations as a means of ensuring that they can effectively manage, understand, and utilize their data. By aligning a company's data observability with its data governance policy, organizations can build a transparent data culture that is protective of organizational assets, compliant with regulatory requirements, and inherently open to collaboration.
May 23, 2023
967 words in the original blog post.
Data reliability engineering focuses on ensuring data quality and reliability in the modern data stack by monitoring, alerting, testing, and documenting data. It differs from other roles such as data quality engineering, data architect, and data engineering. Metaplane is a tool that helps implement best practices for data reliability engineering. The benefits of this approach include improved decision-making, increased trust in data, and faster time to resolution for data issues. Challenges include maintaining consistency across the entire data stack and managing costs and resources.
May 23, 2023
650 words in the original blog post.
The article discusses how to evaluate data observability tools, which are crucial for organizations that rely heavily on data-driven decision making. Data engineering teams often face massive tech debt and ambitious roadmaps, leading to data quality issues being discovered by stakeholders rather than engineers. This erodes trust in the data platform. To address this issue, data observability tools should improve baseline testing and alerting strategies using predictive models and machine learning-based anomaly detection. They should also facilitate test-driven development and provide comprehensive out-of-the-box features. The evaluation of these tools requires broad implementation across all data assets and a trial period of about 30 days to understand their effectiveness fully. Adopting a data observability tool can help diagnose existing quality issues, prevent future ones, and improve overall trust in the data platform.
May 23, 2023
1,166 words in the original blog post.
Data Product Managers are a new role emerging in response to the increasing importance of data engineering and business intelligence in modern companies. They focus on managing the lifecycle of data products, such as business intelligence tools, machine learning models, and other data-driven solutions. Their responsibilities include making critical decisions about data products that impact strategic objectives, collaborating with stakeholders and users, identifying necessary data sources, and spearheading A/B testing of their data products. Key skills required for this role include understanding data warehousing concepts, expertise in machine learning, experience creating data visualizations, and the ability to navigate the DataOps workflow.
May 23, 2023
619 words in the original blog post.
Data contracts are essential for effective data governance, bridging the gap between business and data, and promoting transparency and trust. Clear definitions, metrics, and ownership are critical to the success of data contracts, along with regular updates and reviews. Implementing data contracts requires good communication, collaboration, and persistence. Metaplane provides monitoring and troubleshooting tools to help businesses implement data contracts effectively and improve data quality.
May 23, 2023
1,032 words in the original blog post.
Stale data can lead to poor decision-making in businesses, resulting in lost revenue and unsatisfied customers. Stale data refers to outdated or inaccurate information that may be used for crucial decisions. Sales, marketing, and customer service departments are particularly vulnerable to the negative effects of stale data. Poor data management practices, such as lack of monitoring and data pipeline bugs, can cause stale data. Solutions include implementing data refresh schedules and using data observability tools like Metaplane to ensure data freshness and quality.
May 23, 2023
711 words in the original blog post.
Backfilling Data in 2023 is an essential process for maintaining a complete and accurate historical record of data. It involves populating missing data points into a system to ensure reliability in analysis and decision-making. However, navigating the nuances of backfilling in modern data stacks can be challenging due to their complexity and demand for real-time data processing.
Data observability platforms play a crucial role in mitigating these challenges by offering real-time monitoring and troubleshooting tools. They help identify data drift, pipeline errors, and maintain data quality standards and regulations. Automation, data validation post-backfill, testing in a staging environment, utilization of modern ELT tools, strong data governance principles, parallelizing backfill jobs, handling timezone discrepancies, maintaining open communication, planning for rollback, employing data observability platforms, resource management, and monitoring performance are some best practices to follow while backfilling data in 2023.
In conclusion, businesses must adopt these best practices and leverage the right tools and technologies to maintain data accuracy and completeness, improve decision-making, and succeed in their data efforts.
May 23, 2023
1,483 words in the original blog post.
The future of business intelligence is transforming with the rise of modular data stacks, data observability, embedded analytics, decision intelligence, and collaborative intelligence. These shifts enable faster innovation, higher flexibility, greater scalability, improved trust in data quality, enhanced user experience, smarter decision-making, and increased collaboration across teams and organizations. Metaplane's data observability platform is designed to help data teams overcome challenges in ensuring trust in their data stack.
May 23, 2023
1,199 words in the original blog post.
This article discusses data quality metrics for data warehouses, which are essential for improving the reliability and usefulness of data. It introduces intrinsic and extrinsic data quality dimensions that can be used to measure various aspects of data quality. Intrinsic dimensions include accuracy, completeness, consistency, privacy and security, and freshness, while extrinsic dimensions depend on specific use cases and include relevance, reliability, timeliness, usability, and validity. The article suggests starting from the most important use cases for data in an organization to identify relevant metrics and improve data quality over time using a combination of people, process, and technology strategies.
May 22, 2023
3,689 words in the original blog post.
Data observability is crucial for modern businesses as it helps maintain trust in data by addressing issues such as incorrect information and feature drift over time. Convincing stakeholders to invest in data observability can be challenging, but there are four common reasons that justify this investment: saving engineering time, increasing data team leverage, avoiding costly lapses in data quality, and preserving trust. By implementing a data observability tool, businesses can save time and money while ensuring the accuracy of their data, ultimately leading to better decision-making and improved customer satisfaction.
May 20, 2023
2,387 words in the original blog post.
Monte Carlo data and methods utilize randomness and probabilistic techniques to solve complex real-world problems across various fields, including particle physics, stock price forecasting, and disease outbreak tracking. The origins of Monte Carlo methods can be traced back to the 18th century with Buffon's needle experiment, while significant advancements were made during the Manhattan Project in the late 1940s.
Monte Carlo methods are problem-solving techniques that leverage random sampling and statistical analysis to estimate outcomes and assess uncertainties. They prove particularly valuable when dealing with problems having numerous variables and where analytical solutions are difficult or non-existent. The general Monte Carlo algorithm consists of defining the problem, identifying relevant random variables, generating random samples, evaluating the quantity of interest, and repeating the process to improve accuracy.
Examples of Monte Carlo methods include approximating the value of Pi, estimating the value of an infinite series, and computing integrals. Various tools and libraries are available in different programming languages for conducting Monte Carlo simulations, such as Python's `emcee` package, R's `mc2d` package, and MATLAB's Statistics and Machine Learning Toolbox.
The reliability of Monte Carlo simulations depends on data quality, which directly impacts estimations, sampling, model validity, risk assessment, sensitivity analysis, and generalizability. Effective interpretation of Monte Carlo simulation results involves using summary statistics, visualization, confidence intervals, sensitivity analysis, convergence and error analysis, hypothesis testing, decision-making, and acknowledging the limitations of simulations.
Monte Carlo methods have diverse applications across various domains such as physics, finance, artificial intelligence, gaming, and engineering. Techniques like variance reduction help mitigate statistical uncertainty, while advanced methods like Markov Chain Monte Carlo (MCMC) and Quantum Monte Carlo (QMC) facilitate sampling from complex probability distributions and simulating quantum systems, respectively.
May 20, 2023
1,966 words in the original blog post.
This blog post delves into the differences between DevOps and DataOps methodologies. While both share similarities such as collaboration, automation, and feedback loops, they have distinct focuses. DevOps is centered around application development and deployment, while DataOps aims to optimize the entire data lifecycle from data ingestion to analytics. Both require a cultural shift in the organization and can improve the speed and quality of software development and data management, ultimately driving better business outcomes.
May 18, 2023
1,059 words in the original blog post.
Metaplane introduces tags to help users organize and manage their data more effectively. Users can add tags to dbt models or create custom ones for additional context. Tags are useful not only for organization but also for setting up alert routing based on specific tag criteria. This feature is now available for all Metaplane customers, with a 14-day trial offered for new users.
May 17, 2023
442 words in the original blog post.
dbt (data build tool) is an open-source command-line tool that enables data teams to orchestrate data transformations and check data quality. It features built-in data quality checks, which help maintain data accuracy and consistency across the entire data processing pipeline. Implementing dbt data quality checks involves identifying the data to be tested, defining testing criteria using SQL queries, setting up a dbt project, and running tests manually or on a schedule. Other useful tools for maintaining data quality include Great Expectations, Airflow, and Snowflake.
May 17, 2023
1,136 words in the original blog post.
Machine learning can be used to perform robust data quality checks, offering advantages such as scalability, adaptability to changes in data patterns, and the ability to detect unknown issues. However, it may not always be necessary or ideal for every scenario due to its complexity and overhead. When deploying ML, considerations include incorporating user feedback, evaluating model performance with appropriate metrics, addressing oversampling and imbalanced data, and utilizing open-source tools like Prophet or commercial offerings like Metaplane.
May 16, 2023
1,781 words in the original blog post.
The concept of data observability, which refers to the degree of visibility into our data systems, is a modern adaptation of software observability that draws from control theory. Data observability helps maintain order within our data by preventing lapses in data quality and unacceptable data downtime. It goes beyond data quality monitoring by striving to understand why problems occurred and how they can be prevented. The four pillars of data observability are: metrics, metadata, lineage, and logs. These components are necessary and sufficient to describe the state of a data system at any point in time, in arbitrary detail.
May 15, 2023
2,088 words in the original blog post.
This guide provides an overview of the decision-making process for choosing a data integration tool. It covers important considerations such as data source coverage, extensibility, destination coverage, pricing, security, replication customization, and data transformations. The article also suggests suitable tools for different types of organizations, including early-stage startups, growth-stage startups, data consultancies or marketing agencies, and enterprises.
May 15, 2023
2,544 words in the original blog post.
Data observability is becoming increasingly important as data usage grows across companies regardless of size or industry. Prioritizing data quality and implementing better observability can help maintain trust in data, prevent data loss, utilize historical data effectively, improve decision-making speed, and prioritize tasks more efficiently. Companies should not wait to address these issues but rather adopt data observability platforms or build their own tools proactively.
May 15, 2023
1,537 words in the original blog post.
Data lineage is the process of tracking data through its life cycle across various transformations, and it plays a crucial role in managing and understanding your data landscape. It provides visibility into how data flows and transforms across systems, which is invaluable for data governance, compliance, debugging, and change impact analysis. In Snowflake, you can leverage database metadata to capture data lineage using views like `OBJECT_DEPENDENCIES`, `ACCESS_HISTORY`, and `QUERY_HISTORY`. However, raw data lineage information is most valuable when visualized and integrated into your workflows. Data lineage can help with a number of jobs-to-be-done including root cause analysis, impact analysis, reducing inefficiencies, and onboarding new members of the team or data consumers.
May 15, 2023
5,125 words in the original blog post.
This article discusses four methods to track update times for Google BigQuery tables and views, ensuring data freshness and timely updates. These methods include using the MAX function on a timestamp column, checking metadata's last modified time, leveraging Information Schema's last change time, and utilizing partitioning and clustering features. Each method has its specific use cases and limitations, making it essential to understand these nuances and select the most appropriate one based on individual needs.
May 14, 2023
784 words in the original blog post.
Data freshness is crucial for accurate decision-making, as outdated information can lead to misguided decisions and regulatory non-compliance. Two methods to determine the last update time for Snowflake tables and views are using the MAX function and leveraging the LAST_ALTERED column in system views. The MAX function returns the maximum value of a specified timestamp column, while the LAST_ALTERED column provides system-level information on table or view updates. However, it's essential to consider whether you're dealing with a table or a view and whether you're interested in changes to data or structure when using these methods.
May 14, 2023
758 words in the original blog post.
This article discusses three methods to retrieve row counts for Snowflake tables and views: using the `COUNT(*)` function, leveraging table statistics, and querying metadata or information schema. The first method is straightforward but can be resource-intensive for large tables. The second method provides approximate row counts based on table statistics, while the third method allows efficient retrieval of exact row counts for tables only. Each method has its pros and cons, and users should choose the most suitable one based on their specific requirements.
May 12, 2023
481 words in the original blog post.
Amazon Redshift provides three methods for retrieving row counts in tables and views: using the COUNT function, system statistics, and combining multiple methods. The COUNT function allows an exact count of rows, while system statistics provide approximate counts but are updated automatically. Combining both methods ensures accuracy in tracking row counts over time.
May 12, 2023
496 words in the original blog post.
Google BigQuery offers several methods for retrieving row counts from tables and views, each with its own set of advantages and challenges. The `COUNT(*)` function provides an accurate count but can be resource-intensive for large tables. Using the `INFORMATION_SCHEMA` is efficient and cost-effective but offers only approximate counts that may not reflect recent changes. The BigQuery API allows for programmatic access and consistency in results, though it requires proper authentication and could lead to increased network overhead. Table statistics provide a fast estimate of row counts without full table scans, but these may not be up-to-date after recent modifications. Each method has trade-offs, such as the balance between accuracy, resource consumption, and cost, and the most suitable approach depends on specific needs, such as table size and frequency of updates.
May 12, 2023
862 words in the original blog post.
Data quality tests are processes and procedures implemented to assess the reliability, accuracy, consistency, completeness, and relevance of data within a system or database. These tests help identify potential issues or anomalies that may impact the integrity or usefulness of the data. Some popular ways to apply data quality checks include completeness test, consistency test, accuracy test, integrity test, validity test, deduplication test, timeliness test, uniqueness test, data profiling, and data consistency test. Data quality tests can be implemented anywhere along your ETL process, typically through stored procedures running one-off SQL statements within your data warehouse.
May 10, 2023
1,330 words in the original blog post.
The Four Pillars of Data Observability are metrics, metadata, lineage, and logs. Metrics describe the internal characteristics of data, such as summary statistics or accuracy. Metadata provides external information about data, like volume, structure, and freshness. Lineage outlines dependencies between pieces of data, while logs document interactions between data and the real world, including machine-machine and machine-human interactions. These four pillars together provide a comprehensive understanding of an organization's data health at any given time.
May 10, 2023
2,046 words in the original blog post.
Poor data quality often stems from five fundamental causes: input errors, infrastructure failures, incorrect transformations, invalid assumptions, and ontological misalignment. Understanding these root causes can help businesses improve their data quality management by focusing on what is knowable and controllable. By adopting a systematic approach and learning from other communities, the number and severity of data quality issues can be reduced.
May 03, 2023
1,536 words in the original blog post.
Metaplane introduces an improved way to manage schema change notifications, allowing users to filter out unwanted alerts for specific parts of their data environment. The updated settings include toggles for every database, schema, and table, making it easier to switch off or on notifications as needed. This feature is designed to save time and effort in configuring Metaplane to fit a company's structure and workflows.
May 03, 2023
307 words in the original blog post.