Home / Companies / Gretel.ai / Blog / July 2022

July 2022 Summaries

4 posts from Gretel.ai

Filter
Month: Year:
Post Summaries Back to Blog
Gretel, a platform that helps users access production data, has introduced new product and technology initiatives to ensure its growth and evolution with the needs of modern data consumers. The company's core toolkit includes synthesizing, classifying, and transforming data. To support various synthetic models for different tasks, Gretel has developed the Model Integration Framework (MIF). The platform is also working on automating synthetic data generation to lower the barrier to entry for users.
Jul 26, 2022 1,266 words in the original blog post.
Gretel has introduced the Evaluate API, which generates a Synthetic Data Quality Score (SQS) report for any synthetic dataset by comparing it against real-world data. This allows users to evaluate the quality of existing synthetic data regardless of its source. The Evaluate API is available via CLI and SDK and can be run in the Gretel Cloud or locally to generate a Gretel Synthetic Report.
Jul 20, 2022 774 words in the original blog post.
The article discusses the evaluation of sampling procedures on the quality of synthetic tabular data using Gretel.ai's Synthetic Quality Score (SQS). It explains how to calculate and interpret the SQS, which measures inter-columnar correlations, variance directions via principal component analysis, and discrete mass distributions of each feature in a dataset. The article also explores different sampling methods and their impact on the quality of synthetically generated data using an ensemble of sampling methods that performs as well as direct sampling from the categorical distribution while reducing SQS performance variance.
Jul 13, 2022 720 words in the original blog post.
Data simulation is a process that uses large quantities of data to mimic real-world scenarios or conditions, enabling the creation of comprehensive models of complex systems and facilitating data-driven decision making. It can be used for predicting future events, determining the best course of action, validating AI/ML models, testing hypotheses, understanding relationships, improving predictions, studying phenomena that are difficult to investigate directly, and generating synthetic data representative of specific populations or conditions. Data simulation is highly valuable across various industries and fields of study, with its benefits including enabling the creation of comprehensive models of complex systems, empowering data-driven decision making and strategic planning, helping test hypotheses, understanding relationships, improving predictions, allowing the study of phenomena that are difficult to investigate directly, and generating synthetic data.
Jul 13, 2022 2,409 words in the original blog post.