Home / Companies / Arize / Blog / May 2021

May 2021 Summaries

3 posts from Arize

Filter
Month: Year:
Post Summaries Back to Blog
ML observability refers to tools and practices that help teams monitor and understand the performance of their machine learning (ML) models in real-world scenarios. This is particularly important as more teams adopt ML to streamline their businesses or turn impractical technologies into reality. The challenge lies in translating research lab models to production environments, where data and feature transformations can be inconsistent, leading to poor model performance. By applying evaluation stores and introspection techniques, teams can identify gaps in training data, detect underperforming model slices, compare model performances, validate models, and troubleshoot issues in real-time, ultimately improving their ML efforts.
May 27, 2021 364 words in the original blog post.
Beyond traditional monitoring, observability is a crucial aspect of understanding the health of complex data-driven systems. It enables teams to identify issues such as duplicate or stale data, model drift, and biased training datasets that can lead to unintended consequences. Observability provides granular information about data quality, schema changes, lineage, freshness, distribution, volume, and other key pillars of data health, allowing teams to detect problems early and prevent them from becoming bigger issues. Unlike monitoring, observability enables active learning, root cause analysis, and collaboration across cross-functional teams to resolve data issues before they impact the business. By applying principles of software application observability and reliability to data and ML, teams can build more trustworthy and reliable systems, gain insight into model performance, detect drift, and identify the "why" behind broken data pipelines and failed models. Ultimately, observability is essential for building a culture of trust in data-driven systems and making informed decisions based on accurate insights.
May 19, 2021 1,562 words in the original blog post.
The concept of data ethics is being reevaluated by a group of AI researchers in Africa, who argue that the origin, collection, and sharing of data are often overlooked but critical components of AI ethics. They highlight issues such as deficit narratives, extractive data practices, and moral distance between data collectors and communities, which can lead to harm and mistrust. The authors emphasize the importance of building trust, respecting local norms and contexts, and ensuring that data is shared in a way that benefits both communities and science. They also argue that AI ethics must start with data collection and sharing, rather than just model development, to ensure fairness, equity, and justice in AI systems.
May 06, 2021 1,297 words in the original blog post.