{"id":144941,"date":"2026-08-19T14:24:14","date_gmt":"2026-08-19T14:24:14","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/144941\/"},"modified":"2026-08-19T14:24:14","modified_gmt":"2026-08-19T14:24:14","slug":"next-generation-synthetic-trials-in-hematology-with-generative-artificial-intelligence","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/144941\/","title":{"rendered":"Next-generation synthetic trials in hematology with generative artificial intelligence"},"content":{"rendered":"<p>Despite their potential, synthetic patient data are no panacea, requiring clinicians and researchers to be aware of potential pitfalls and adhere to best practices for implementation (Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41375-026-03116-9#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>). First, data generation itself is locked in a conundrum: Synthetic data may enable privacy-compliant health data access and sharing as well as facilitate novel trial designs, yet to train a generative model, one needs a large, diverse, and representative training sample so that the model can infer accurate feature distributions and capture intricate relationships. Generative models cannot make up useful data from scratch. If one desires to generate synthetic data of a specific group of patients, yet one has no access to a sufficiently sized training cohort in the first place, no generative model will be able to create said group of patients from thin air. Hence, the first step is defining the relevant task to be addressed with synthetic patients and then identifying a group of patients that can actually be synthetically created. For clinical trials, this may best apply to patients receiving standard of care since plenty of patient records for training exist across healthcare systems, previous trials, and registries.<\/p>\n<p>Fig. 2: Pitfalls and the path forward for synthetic data in clinical trials.<img decoding=\"async\" aria-describedby=\"figure-2-desc\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/41375_2026_3116_Fig2_HTML.png\" alt=\"Fig. 2: Pitfalls and the path forward for synthetic data in clinical trials.\" loading=\"lazy\" width=\"685\" height=\"361\"\/><\/p>\n<p>Off-the-shelf synthetic data are no silver bullet: Synthetic data could amplify existing biases in medicine, lead to further underrepresentation of minority groups, leak sensitive patient information, and allow for adversarial behavior. Hence, following best practices is essential, including multi-source training, deliberate inclusion of minority populations, privacy audits and safeguards, and transparent disclosure of training data properties and model settings. In the setting of clinical trials, commercial or academic conflicts of interest may require synthetic data generation by third parties to mitigate cherry-picking and uphold regulatory standards. Created with Microsoft PowerPoint.<\/p>\n<p>Second, generative models mirror feature distributions of their training data. Hence, they may also carry over or even amplify implicit biases, including local patient demographics, institutional preferences in treatment selection, or specific properties of subgroups if they are overrepresented. This is further complicated by inadvertently modeling confounding covariates or neglecting unknown or unknowable covariates modifying statistical properties. Training data should therefore come from multiple sources &#8211; ideally not only relying on patient data from previously conducted clinical trials but also include real-world registries [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 14\" title=\"Passamonti F, Corrao G, Castellani G, Mora B, Maggioni G, Della Porta MG, et al. Using real-world evidence in haematology. Best Pr Res Clin Haematol. 2024;37:101536.\" href=\"http:\/\/www.nature.com\/articles\/s41375-026-03116-9#ref-CR14\" id=\"ref-link-section-d890915e507\" rel=\"nofollow noopener\" target=\"_blank\">14<\/a>, <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Gale RP, Hochhaus A. Are real- world data real-world data? Leukemia. 2025;39:2311&#x2013;2.\" href=\"http:\/\/www.nature.com\/articles\/s41375-026-03116-9#ref-CR15\" id=\"ref-link-section-d890915e510\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>] &#8211; to ensure adequate sample size, representativeness, and generalizability [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 16\" title=\"Passamonti F, Corrao G, Castellani G, Mora B, Maggioni G, Gale RP, et al. The future of research in hematology: integration of conventional studies with real-world data and artificial intelligence. Blood Rev. 2022;54:100914.\" href=\"http:\/\/www.nature.com\/articles\/s41375-026-03116-9#ref-CR16\" id=\"ref-link-section-d890915e513\" rel=\"nofollow noopener\" target=\"_blank\">16<\/a>]. All variables should be assessed for interference or redundancies prior to model training. The inclusion of minority groups is of particular importance, as otherwise generative models will not contribute to solving the equity problem in medicine but will quietly intensify it.<\/p>\n<p>Third, by modeling the underlying distribution of features and creating synthetic patient samples that fit this distribution, synthetic data are not exact replicas of the training data patients. Yet, privacy preservation is not guaranteed by default. Sensitive information can still be exposed either through unintended model behavior or adversarial manipulation, such as membership inference or model inversion attacks [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 17\" title=\"Hu H, Salcic Z, Sun L, Dobbie G, Yu PS, Zhang X. Membership inference attacks on machine learning: a survey. ACM Comput Surv. 2022;54:235:1&#x2013;235:37.\" href=\"http:\/\/www.nature.com\/articles\/s41375-026-03116-9#ref-CR17\" id=\"ref-link-section-d890915e519\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>, <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 18\" title=\"Yang W, Wang S, Wu D, Cai T, Zhu Y, Wei S, et al. Deep learning model inversion attacks and defenses: a comprehensive survey. Artif Intell Rev. 2025;58:242.\" href=\"http:\/\/www.nature.com\/articles\/s41375-026-03116-9#ref-CR18\" id=\"ref-link-section-d890915e522\" rel=\"nofollow noopener\" target=\"_blank\">18<\/a>]. Safeguarding patients\u2019 privacy should not be an afterthought, but privacy audits should rather guide the data generation process from inception. Crucially, one must account for a privacy-usability tradeoff: The more synthetic data differ from original data, the more privacy-compliant they become, yet the less useful they are for downstream tasks. In the absence of universally accepted thresholds for adequate privacy preservation, potential information leakage and downstream usability must be iteratively assessed throughout generation. Differential privacy budgets, real-to-synthetic distance metrics, and design safeguards combatting adversarial attacks help mitigate the risk of privacy breaches [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 19\" title=\"Dwork, C. Differential Privacy. In: Bugliesi M, Preneel B, Sassone V, Wegener I (eds). Automata, Languages and Programming. Springer: Berlin, Heidelberg, 2006, pp 1&#x2013;12.\" href=\"http:\/\/www.nature.com\/articles\/s41375-026-03116-9#ref-CR19\" id=\"ref-link-section-d890915e525\" rel=\"nofollow noopener\" target=\"_blank\">19<\/a>, <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 20\" title=\"Pilgram L, Dankar FK, Drechsler J, Elliot M, Domingo-Ferrer J, Francis P, et al. A consensus privacy metrics framework for synthetic data. Patterns. 2025;6:101320.\" href=\"http:\/\/www.nature.com\/articles\/s41375-026-03116-9#ref-CR20\" id=\"ref-link-section-d890915e528\" rel=\"nofollow noopener\" target=\"_blank\">20<\/a>].<\/p>\n<p>Lastly, synthetic data are in a regulatory limbo. Regulatory agencies increasingly acknowledge the need for alternative control cohorts in settings where placebo control is not feasible, or recruitment is slowed by small or inaccessible patient populations, especially in rare cancers or molecular subgroups [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 21\" title=\"Serrano C, Rothschild S, Villacampa G, Heinrich MC, George S, Blay J-Y, et al. Rethinking placebos: embracing synthetic control arms in clinical trials for rare tumors. Nat Med. 2023;29:2689&#x2013;92.\" href=\"http:\/\/www.nature.com\/articles\/s41375-026-03116-9#ref-CR21\" id=\"ref-link-section-d890915e534\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a>, <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"Farah E, Kenney M, Warkentin MT, Cheung WY, Brenner DR. Examining external control arms in oncology: A scoping review of applications to date. Cancer Med. 2024;13:e7447.\" href=\"http:\/\/www.nature.com\/articles\/s41375-026-03116-9#ref-CR22\" id=\"ref-link-section-d890915e537\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a>]. However, currently no regulatory framework exists for synthetically controlled trials. Regulators will have to define appropriate quality measures, similar to existing frameworks such as the Food and Drug Administration\u2019s \u201cGood Machine Learning Practice for Medical Device Development\u201d[<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 23\" title=\"Food and Drug Administration. Good Machine Learning Practice for Medical Device Development: Guiding Principles. FDA 2025. &#010;                https:\/\/www.fda.gov\/medical-devices\/software-medical-device-samd\/good-machine-learning-practice-medical-device-development-guiding-principles&#010;                &#010;               (accessed 5 Feb2025).\" href=\"http:\/\/www.nature.com\/articles\/s41375-026-03116-9#ref-CR23\" id=\"ref-link-section-d890915e540\" rel=\"nofollow noopener\" target=\"_blank\">23<\/a>]. These should include transparency requirements on training cohort properties, potentially arising limitations and biases, disclosure of model architecture, metrics for fidelity and usability, as well as privacy preservation [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 24\" title=\"van Breugel B, Liu T, Oglic D, van der Schaar M. Synthetic data in biomedicine via generative artificial intelligence. Nat Rev Bioeng. 2024;2:991&#x2013;1004.\" href=\"http:\/\/www.nature.com\/articles\/s41375-026-03116-9#ref-CR24\" id=\"ref-link-section-d890915e543\" rel=\"nofollow noopener\" target=\"_blank\">24<\/a>, <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 25\" title=\"Eckardt J-N, Hahn W, Prelaj A, Bornh&#xE4;user M, Middeke JM, Kather JN. Artificial intelligence-generated synthetic data for cancer research and clinical trials. Nat Rev Cancer. 2026;26:351&#x2013;63.\" href=\"http:\/\/www.nature.com\/articles\/s41375-026-03116-9#ref-CR25\" id=\"ref-link-section-d890915e546\" rel=\"nofollow noopener\" target=\"_blank\">25<\/a>]. Crucially, synthetic data generation allows a degree of customization, e.g., by generating large cohorts and selecting only those cases with the desired properties. In clinical trials, this may lead to a critical conflict of interest as entities with commercial or academic stakes in the outcome of a trial could potentially \u201ccherry-pick\u201d a synthetic control cohort to manufacture a desired result. Regulatory agencies should therefore provide a framework for synthetic data generation by independent third parties, with the resulting cohort withheld from investigators and sponsors until completion of intervention arm data collection.<\/p>\n<p>In summary, synthetic data hold the potential to reduce barriers in data sharing and may enable novel trial designs, accelerate recruitment, reduce failure rates, and provide more patients with access to investigational therapies, particularly in the era of precision therapies in hematology. Yet, they are no silver bullet: Rigorous evaluation, quality assessment, privacy preservation, and regulatory guidance are needed before clinical implementation.<\/p>\n","protected":false},"excerpt":{"rendered":"Despite their potential, synthetic patient data are no panacea, requiring clinicians and researchers to be aware of potential&hellip;\n","protected":false},"author":2,"featured_media":144942,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,25,7044,375,617,71362,50874,50873,5368,21053,5813],"class_list":["post-144941","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-artificial-intelligence","tag-cancer-research","tag-clinical-trials","tag-general","tag-haematological-cancer","tag-hematology","tag-intensive-critical-care-medicine","tag-internal-medicine","tag-medicine-public-health","tag-oncology"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/144941","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=144941"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/144941\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/144942"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=144941"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=144941"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=144941"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}