Avoiding Data Leakage in Synthetic Data Projects
Synthetic data offers vast opportunities for machine learning model training without real-world constraints. Yet, beneath this promise lies the risk of…
Read more →Gretel is a synthetic data platform built to help teams create artificial datasets that preserve the statistical shape of real data without exposing the underlying sensitive records. Founded in 2019 and later acquired by NVIDIA, its tools let developers generate, anonymize, and share structured and text data through APIs and SDKs rather than moving real production data around.
The platform focuses on generating synthetic tabular and text data, applying techniques such as differential privacy and offering automated scoring that reports both how useful the synthetic data is and how well it protects privacy. This lets teams stand up realistic datasets for model training, testing, and sharing across teams or organizations while reducing exposure of regulated information.
Gretel sits at the synthetic-data layer, an alternative or complement to sourcing real data when privacy, scarcity, or class imbalance make real datasets hard to use. Rather than collecting data from the web or vertical sources, it manufactures data that mirrors real distributions, which is increasingly relevant as AI labs exhaust readily available real-world data. Under NVIDIA, its capabilities are positioned as part of a broader cloud-based toolkit for training and developing AI models.
TabularText