Despite their potential, synthetic patient data are no panacea, requiring clinicians and researchers to be aware of potential pitfalls and adhere to best practices for implementation (Fig. 2). First, data generation itself is locked in a conundrum: Synthetic data may enable privacy-compliant health data access and sharing as well as facilitate novel trial designs, yet to train a generative model, one needs a large, diverse, and representative training sample so that the model can infer accurate feature distributions and capture intricate relationships. Generative models cannot make up useful data from scratch. If one desires to generate synthetic data of a specific group of patients, yet one has no access to a sufficiently sized training cohort in the first place, no generative model will be able to create said group of patients from thin air. Hence, the first step is defining the relevant task to be addressed with synthetic patients and then identifying a group of patients that can actually be synthetically created. For clinical trials, this may best apply to patients receiving standard of care since plenty of patient records for training exist across healthcare systems, previous trials, and registries.

Fig. 2: Pitfalls and the path forward for synthetic data in clinical trials.Fig. 2: Pitfalls and the path forward for synthetic data in clinical trials.

Off-the-shelf synthetic data are no silver bullet: Synthetic data could amplify existing biases in medicine, lead to further underrepresentation of minority groups, leak sensitive patient information, and allow for adversarial behavior. Hence, following best practices is essential, including multi-source training, deliberate inclusion of minority populations, privacy audits and safeguards, and transparent disclosure of training data properties and model settings. In the setting of clinical trials, commercial or academic conflicts of interest may require synthetic data generation by third parties to mitigate cherry-picking and uphold regulatory standards. Created with Microsoft PowerPoint.

Second, generative models mirror feature distributions of their training data. Hence, they may also carry over or even amplify implicit biases, including local patient demographics, institutional preferences in treatment selection, or specific properties of subgroups if they are overrepresented. This is further complicated by inadvertently modeling confounding covariates or neglecting unknown or unknowable covariates modifying statistical properties. Training data should therefore come from multiple sources – ideally not only relying on patient data from previously conducted clinical trials but also include real-world registries [14, 15] – to ensure adequate sample size, representativeness, and generalizability [16]. All variables should be assessed for interference or redundancies prior to model training. The inclusion of minority groups is of particular importance, as otherwise generative models will not contribute to solving the equity problem in medicine but will quietly intensify it.

Third, by modeling the underlying distribution of features and creating synthetic patient samples that fit this distribution, synthetic data are not exact replicas of the training data patients. Yet, privacy preservation is not guaranteed by default. Sensitive information can still be exposed either through unintended model behavior or adversarial manipulation, such as membership inference or model inversion attacks [17, 18]. Safeguarding patients’ privacy should not be an afterthought, but privacy audits should rather guide the data generation process from inception. Crucially, one must account for a privacy-usability tradeoff: The more synthetic data differ from original data, the more privacy-compliant they become, yet the less useful they are for downstream tasks. In the absence of universally accepted thresholds for adequate privacy preservation, potential information leakage and downstream usability must be iteratively assessed throughout generation. Differential privacy budgets, real-to-synthetic distance metrics, and design safeguards combatting adversarial attacks help mitigate the risk of privacy breaches [19, 20].

Lastly, synthetic data are in a regulatory limbo. Regulatory agencies increasingly acknowledge the need for alternative control cohorts in settings where placebo control is not feasible, or recruitment is slowed by small or inaccessible patient populations, especially in rare cancers or molecular subgroups [21, 22]. However, currently no regulatory framework exists for synthetically controlled trials. Regulators will have to define appropriate quality measures, similar to existing frameworks such as the Food and Drug Administration’s “Good Machine Learning Practice for Medical Device Development”[23]. These should include transparency requirements on training cohort properties, potentially arising limitations and biases, disclosure of model architecture, metrics for fidelity and usability, as well as privacy preservation [24, 25]. Crucially, synthetic data generation allows a degree of customization, e.g., by generating large cohorts and selecting only those cases with the desired properties. In clinical trials, this may lead to a critical conflict of interest as entities with commercial or academic stakes in the outcome of a trial could potentially “cherry-pick” a synthetic control cohort to manufacture a desired result. Regulatory agencies should therefore provide a framework for synthetic data generation by independent third parties, with the resulting cohort withheld from investigators and sponsors until completion of intervention arm data collection.

In summary, synthetic data hold the potential to reduce barriers in data sharing and may enable novel trial designs, accelerate recruitment, reduce failure rates, and provide more patients with access to investigational therapies, particularly in the era of precision therapies in hematology. Yet, they are no silver bullet: Rigorous evaluation, quality assessment, privacy preservation, and regulatory guidance are needed before clinical implementation.