Synthetic Data: Filling Gaps While Creating New Uncertainties

Have you ever heard about synthetic data? Image: 3844328, on Pixabay.

 

By Mariana Meneses

As artificial intelligence demands ever larger and more varied datasets, synthetic data has emerged as a powerful way to fill gaps, protect privacy, and simulate rare or difficult scenarios. Its rapid expansion, however, raises a question: how can we trust data designed to resemble reality when the differences may be difficult to detect, yet consequential for science, medicine and society? 

According to Cognilytica, the market for synthetic data generation was roughly $110 million in 2021. The research firm expects that to reach $1.15 billion by 2027. Grand View Research anticipates the AI training dataset market to reach more than $8.6 billion by 2030, representing a compound annual growth rate (CAGR) of just over 22%.” – Taryn Plumb/Venturebeat (2022) 

In January 2025, The Guardian’s Global Technology Editor Dan Milmo reported that Elon Musk believed AI companies had largely exhausted the available human-generated data used for training and would increasingly turn to synthetic data produced by AI systems themselves. Musk noted that this approach was already being used across the industry but warned that AI hallucinations could introduce unreliable information into training material. Andrew Duncan, from the Alan Turing Institute, cautioned that excessive reliance on synthetic data could lead to diminishing returns and “model collapse,” while the growing amount of AI-generated content online may make it harder to prevent synthetic material from re-entering future training datasets. 

This is felt particularly in industries like healthcare and finance, where tight privacy regulations are making the shortage of data even more acute.” – Gediminas Rickevicius/Dataversity (2025) 

What is Synthetic Data?

In “3 Questions: The pros and cons of synthetic data in AI”, published in in September 2025, MIT News, writer Adam Zewe explains that synthetic data are generated by algorithms to reproduce the statistical patterns of real data without directly containing real-world records. They can be created for language, images, audio, and tabular data, often by training generative models on a limited amount of real information. Their use is growing because they can be produced quickly, shared more easily, and adapted to specific organizational needs.

 

“Modern AI systems are rooted in machine learning and deep learning techniques. These methodologies are notorious for their computational intensity, involving complex mathematical processes and algorithms. During the training phase, AI models process large volumes of data, while continuously adapting and refining their parameters to optimize performance, rendering the training process computationally intensive.” In the image, a plot shows the computation used to train notable artificial intelligence systems. Credit: Our World in Data.

 

Zewe explains that the main applications include software testing and machine-learning training. Organizations can generate large volumes of artificial transactions, simulate particular user groups, or supplement rare events, such as financial fraud, when real examples are limited. Moreover, synthetic data can also reduce dependence on sensitive information, lower data-collection costs, and help improve models trained on small datasets. 

In the article entitled “Governing synthetic data in medical research: the time is now”, published in the journal The Lancet Digital Health in April 2025, Daniela Boraschi, from the Kavli Centre for Ethics, Science, and the Public, at the University of Cambridge, and co-authors, argue that synthetic medical data, including text, images and structured records, are already being used in medicine. They note that these datasets can preserve key features of real patient information while being anonymized and adapted for research, potentially supporting more open, efficient and equitable studies in areas such as cancer, cardiovascular disease and COVID-19. 

Synthetic data generated by AI can support scientific research by filling gaps, correcting biases, modelling complex phenomena, protecting privacy, and reducing reliance on human or animal subjects. This is what David B. Resnik, from the US National Institute of Environmental Health Sciences, and co-authors argue in the article “GenAI synthetic data create ethical challenges for scientists. Here’s how to address them”, published in the journal Proceedings of the National Academy of Sciences, in February 2025.  

The United Kingdom’s Guidance on AI Insights: Synthetic Data (HTML), last updated in March 2026, explains that synthetic data are artificially created to reproduce the statistical properties, relationships and distributions of real-world datasets. According to the publication, they are used mainly to improve data quality rather than simply increase quantity, helping represent rare events, rebalance underrepresented groups and reduce privacy risks when sensitive or identifiable information cannot be safely used. These datasets can be generated through methods such as Monte Carlo simulations (also known as multiple probability simulations), normal-distribution sampling and more specialized algorithms for complex data.

 

For more on model collapse, see our April 2025 article Celebrating Human Creativity Means Guarding Against Stereotyping by AI, from March 2024. Image: Gerd Altmann, on Pixabay.

 

Simulation-to-Reality Gap?

In the article “Synthetic data, synthetic trust: navigating data challenges in the digital revolution”, published in the journal The Lancet Digital Health, in November 2025, Arman Koul, from the School of Medicine at Stanford University, and co-authors explain that, although synthetic data can help address limited access to health-care information, protect privacy and augment small datasets, more data does not automatically produce better medical AI.  

Koul and co-authors highlight that artificially generated records may reproduce biases, omit rare diseases and unusual treatment responses, distort relationships among clinical and demographic variables, and create statistically plausible patterns that do not reflect medical reality. This can produce “synthetic trust”: unwarranted confidence in models that appear comprehensive but fail to preserve clinical validity or population diversity.  

The UK government document also warns that synthetic data can reproduce errors, biases and hidden dependencies from the source material or introduce unrealistic artificial patterns. Attempts to remove sensitive features or correct imbalances may also create new distortions if relationships within the data are not carefully understood. 

This is in line with what Assistant Professor Dani Shanley, from the Philosophy Department at Maastricht University, explains in “Synthetic data, real harm”, published in the Ada Lovelace Institute’s Blog also in September 2025. While real-world data are scarce, biased, expensive or sensitive, synthetic data can expand limited datasets, improve representation, and protect privacy. However, it also creates a “simulation-to-reality gap”, as artificial datasets may appear statistically convincing while missing rare cases, complex relationships and important features of the real world that may be unknown to the developers. When synthetic information repeatedly enters AI training pipelines, it can also contribute to feedback loops, data pollution, and model collapse.

 

“Not enough real-world source data” was the leading challenge appointed by respondents in a survey of 150 IT and Data and Analytics leaders who work with or oversee groups that work with AI-generated synthetic data at their organization, conducted in April 2023 by the Gartner Peer Community. Image: Generative AI for Synthetic Data, by the Gartner Peer Community (pdf).

 

Professor Shanley explains that the proposed benefits of synthetic data for fairness and privacy therefore involve difficult trade-offs. Developers may generate examples of underrepresented groups or remove discriminatory patterns, but decisions about what counts as fair representation are inherently value-laden and can reproduce developers’ assumptions and blind spots. Likewise, highly realistic synthetic data may preserve enough detail to reveal or reconstruct private information, meaning that greater usefulness can sometimes reduce privacy protection. 

A similar position is taken in the World Economic Forum’s publication entitled “As AI blurs the lines between real and synthetic data, strong governance is essential”, published in October 2025 and co-authored by Lauren Woodman, CEO from DataKind, a global nonprofit organization connecting AI and data experts with social-impact organizations, and Arun Sundararajan, Professor at NYU Stern School of Business. The article is part of the Annual Meetings of the Global Future Councils and Cybersecurity. 

According to the WEF article, synthetic data has evolved from a niche solution into an important resource for AI development. It can address missing or underrepresented information, protect sensitive data, and create controlled environments for testing scenarios in areas such as healthcare, autonomous vehicles, finance, climate modelling, and infrastructure planning. However, the WEF alerts that, as synthetic data becomes more realistic and widespread, the distinction between artificial and authentic information is becoming harder to maintain. Moreover, biased source data can reproduce or amplify inequalities, while repeatedly training AI systems on AI-generated material can reduce their accuracy and reliability. Synthetic media, including deepfakes and cloned voices, can also weaken public trust in the authenticity of what people see, hear, and read. 

Koul and co-authors note that these risks are particularly serious when large synthetic datasets are generated from very small or unrepresentative samples. Artificially expanding such data can conceal exclusions, amplify inequalities, and reduce model performance for underrepresented groups. Repeatedly training systems on synthetic information may also cause model collapse, gradually reducing diversity and accuracy, while highly realistic synthetic records can still expose sensitive patient characteristics.  

Will AI understand human patterns, or will AI repeatedly shape our patterns?

Resnik and Mohammad Hosseini (2024) note that AI in general introduces risks involving bias, random errors, opacity, and privacy. Because AI processes information differently from humans, it may cause problems by identifying misleading patterns, fabricatingplausible references, or producing results that are difficult even for specialists to explain. Moreover, as Resnik and co-authors (2025) explain, because generative AI can produce highly realistic datasets, images and text, traditional methods for detecting scientific fraud may become less effective.  

The authors warn that poorly disclosed synthetic data could corrupt the scientific record, reduce reproducibility, and even degrade future AI models trained on contaminated information. Problems like synthetic data being used to expose private information through reconstruction, or the amplification of biases contained in their source data, are particularly serious when synthetic information affects regulatory, medical or other high-stakes decisions.

 

Trusting Synthetic Data

Resnik and co-authors (2025) argue that generative AI’s rapid adoption of these technologies has outpaced regulation. There is still no agreement on how synthetic medical data should be generated, evaluated or compared, and many AI processes remain difficult to observe. This raises questions about data quality, generalizability, accountability, and whether findings based on synthetic information accurately represent patients and the realities of health care. The authors call for a governance framework specifically designed for synthetic medical data, supported by robust standards and certification. Developing it will require collaboration among medical professionals, ethicists, machine-learning researchers, and the public to establish shared definitions of quality, safety, fairness, and appropriate use. They argue that this work is urgent if synthetic data are to be integrated responsibly into medical research and practice.  

Professor Shanley similarly urges that synthetic data requires governance and quality assurance designed specifically for its distinctive risks. Shanley believes that this should include transparent documentation, provenance tracking, independent audits of both datasets and generation algorithms, real-world validation, ethical training, and public participation in decisions about representation. According to Shanley, synthetic data are most defensible as a tool for augmentation, experimentation and stress-testing, not as a substitute for reality, and its use should always be assessed in relation to the social or political problem it is supposed to solve. 

To Adam Zewe, synthetic data must be evaluated within the specific system and task for which they are used. Zewe emphasizes task-specific efficacy measures, privacy and quality checks, balanced sampling and systematic evaluation to ensure that synthetic data improve performance rather than introduce new errors. 

Koul and co-authors call for a shift from data quantity to verifiable quality throughout the AI lifecycle. Proposed safeguards include documenting how synthetic data were generated, disclosing biases and gaps, testing clinical plausibility with health professionals, comparing performance against real-world data, tracking data provenance and reidentification risks, and monitoring whether predictions rely too heavily on synthetic features. Continuous real-world validation is essential to ensure that synthetic data preservesthe complexity and diversity of patient populations without compromising fairness, safety or trust. 

The WEF’s publication argues that business leaders should treat synthetic data governance as a distinct strategic priority. Organizations need transparent standards, collaboration among developers, users, executives, lawyers and policymakers, and safeguards such as watermarking and dataset labels. Importantly, they should invest early in traceability and data-provenance systems that record how and when synthetic information enters a dataset, strengthening accountability, and reducing risks. 

The UK’s Guidance emphasizes that synthetic data should complement, not replace, real-world information. It must undergo the same rigorous testing, fairness evaluation and validation as other data, including comparisons with independent real-world datasets. Organizations should also document and version synthetic datasets so they can be examined and reproduced, ensuring that models do not perform well in artificial testing but fail when deployed in real conditions. 

However, evaluating synthetic data is not a purely technical exercise. This is what Louis Ravn, from the University of Amsterdam, and co-authors argue in “Unraveling the Regimes of Synthetic Data Metrics: Expectations, Ethics, and Politics”, published in the Springer Nature Link’s journal Digital Society (2025). Metrics such as utility, privacy, fidelity and fairness shape which datasets are considered safe and legitimate, but in practice usefulness and similarity to real data often receive more attention than privacy, fairness and environmental impacts. This can reduce complex ethical and political questions to numerical scores, create false confidence, and overlook risks such as identifying individuals through patterns retained from the original data. The authors conclude that improving metrics alone is insufficient; evaluation must also address the wider social, ethical, environmental and political consequences that standardized measures may leave out. 

Transparency, Fairness, Reproducibility and Accountability

Resnik and Hosseini (2024) argue that researchers remain accountable for all AI-assisted work and must verify outputs, address bias, disclose limitations, protect confidential information, and clearly label synthetic data. AI systems may contribute to research but should not be credited as authors or inventors because they cannot assume moral or legal responsibility. The authors call for updated guidance that applies established principles such as transparency, fairness, reproducibility and accountability to AI use, supported by accessible explanations, community engagement, ethics training and regularly revised institutional policies. 

Resnik and co-authors (2025) call for clear definitions distinguishing synthetic from real data, based partly on their provenance, or connection to actual phenomena. Researchers should disclose how and why synthetic data were used, identify which information is artificial, and share the relevant datasets, algorithms, and code. The authors argue that scientific institutions should also provide ethical training and develop technical safeguards such as watermarking, certification, detection systems and privacy protections, while recognizing that responsible use ultimately depends on the integrity of researchers.

 

Synthetic data should be a complement to real-world evidence and be validated continuously under real conditions. Image: Gerd Altmann, on Pixabay.

 

Synthetic data can help researchers and organizations fill data gaps, represent rare cases, protect sensitive information, and create controlled environments for testing. Its growing use across medicine, finance, climate modelling, and software development reflects these advantages. At the same time, synthetic data are constructed representations shaped by their source material, generation methods and design choices. They may reproduce bias, omit important cases, introduce artificial relationships, expose private information, and weaken model accuracy when repeatedly reused. 

Responsible use requires treating synthetic data as a complement to real-world evidence and validating it continuously under real conditions.

Clear labelling, documentation, provenance tracking, privacy testing, independent auditing and disclosure of generation methods are essential, alongside accountability from researchers and organizations. Governance must also address the ethical and political judgments involved in defining quality, fairness, privacy, and acceptable risk. The value of synthetic data depends on whether it can support reliable knowledge while remaining transparent about how it was created and where its limitations lie.


Craving more information? Check out these recommended TQR articles:

Enjoyed this? Help us improve.

☞ complete our Short survey

 

Have we made any errors?

Spotted an error or want to contribute your expertise? We’d love to hear from you — reach us at info@thequantumrecord.com. The Quantum Record exists to bring researchers and curious minds together around science and technology that matters.

Leave a Reply

Your email address will not be published. Required fields are marked *

The Quantum Record is a non-profit journal of philosophy, science, technology, and time. The potential of the future is in the human mind and heart, and in the common ground that we all share on the road to tomorrow. Promoting reflection, discussion, and imagination, The Quantum Record highlights the good work of good people and aims to join many perspectives in shaping the best possible time to come. We would love to stay in touch with you, and add your voice to the dialogue.

Join Our Community