The Office-Home Dataset (cited by 600+ papers) contains hundreds of incorrect labels and outliers.
Blog post from Cleanlab
The Office-Home Dataset is a widely used computer vision dataset that contains hundreds of erroneous labels and data issues, including mislabeled examples, ambiguous examples, and outliers, which can be detrimental to modeling and analytics efforts. These errors were discovered using Cleanlab Studio, an automated solution that identifies and fixes data issues using AI. The dataset was curated by collecting images from a web crawler and filtering them to ensure the desired object was in the picture, but this method often produces incorrect image-label pairs. By running the dataset through Cleanlab Studio, researchers can identify and correct these errors, which can improve the accuracy of their models and conclusions.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.