Generating Synthetic Data with AI and Machine Learning

Saturday 22 March 2025


Synthetic data has become a crucial tool in many fields, from medicine to finance, as it allows researchers to test and validate their theories without putting real people or organizations at risk. But creating synthetic data that accurately reflects the characteristics of real-world data can be a daunting task.


Enter the world of artificial intelligence (AI) and machine learning, where researchers have been working on developing algorithms that can generate synthetic data that is not only realistic but also private. This means that the data cannot be traced back to its original source, ensuring that individuals’ privacy is protected.


One such algorithm is called Private Evolution (PE), which uses a combination of AI and machine learning techniques to generate synthetic data that meets high standards of accuracy and privacy. PE works by using a foundation model, which is a type of AI designed to learn from large datasets, to generate random samples. These samples are then iteratively selected based on their similarity to the original private data.


The team behind PE has taken this algorithm to the next level by experimenting with different simulators, rather than just relying on foundation models. Simulators are computer programs that can generate realistic images or text, and they have been used in various fields such as computer graphics and video games.


In their study, the researchers tested PE using three different simulators: a text rendering program, a computer graphics-based render, and a rule-based avatar generator. They found that by combining these simulators with the foundation model, they could generate synthetic data that was not only highly accurate but also private.


The implications of this research are significant. For example, in medicine, PE could be used to create synthetic patient data for testing new treatments or medical devices without compromising patients’ privacy. In finance, it could be used to generate synthetic financial data for training and testing machine learning models without putting individuals’ financial information at risk.


But how does it work? Well, the algorithm starts by using a foundation model to generate random samples. These samples are then iteratively selected based on their similarity to the original private data. The simulator is used to create realistic images or text that reflect the characteristics of the real-world data.


For example, if you’re trying to generate synthetic medical patient data, the simulator might be used to create realistic images of patients with different conditions or treatments. The foundation model would then select these samples based on their similarity to the original private data, ensuring that the generated data is not only realistic but also accurate.


Cite this article: “Generating Synthetic Data with AI and Machine Learning”, The Science Archive, 2025.


Artificial Intelligence, Machine Learning, Synthetic Data, Privacy, Accuracy, Simulators, Computer Graphics, Medical Patient Data, Financial Data, Foundation Model.


Reference: Zinan Lin, Tadas Baltrusaitis, Sergey Yekhanin, “Differentially Private Synthetic Data via APIs 3: Using Simulators Instead of Foundation Model” (2025).


Leave a Reply