{"id":1048283,"date":"2026-06-25T02:34:22","date_gmt":"2026-06-25T02:34:22","guid":{"rendered":"https:\/\/www.europesays.com\/uk\/1048283\/"},"modified":"2026-06-25T02:34:22","modified_gmt":"2026-06-25T02:34:22","slug":"disparate-privacy-risks-from-medical-ai","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/uk\/1048283\/","title":{"rendered":"Disparate privacy risks from medical AI"},"content":{"rendered":"<p>Medical artificial intelligence (AI) has immense potential to improve health outcomes, particularly in regions in which specialized medical expertise is scarce<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 1\" title=\"Fleming, K. A. et al. The Lancet Commission on diagnostics: transforming access to diagnostics. Lancet 398, 1997&#x2013;2050 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR1\" id=\"ref-link-section-d97577649e462\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>. At the same time, AI also poses new challenges and risks, including security vulnerabilities that arise when models are deployed. Untrusted users with access to an AI model may, by merely observing its predictions, steal its parameters<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 8\" title=\"Tram&#xE8;r, F., Zhang, F., Juels, A., Reiter, M. K. &amp; Ristenpart, T. Stealing machine learning models via prediction APIs. In Proc. 25th USENIX Security Symposium 601&#x2013;618 (USENIX, 2016).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR8\" id=\"ref-link-section-d97577649e466\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 9\" title=\"Carlini, N. et al. Stealing part of a production language model. In Proc. 41st International Conference on Machine Learning 5680&#x2013;5705 (ICML, 2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR9\" id=\"ref-link-section-d97577649e469\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a> or perform privacy attacks<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Shokri, R., Stronati, M., Song, C. &amp; Shmatikov, V. Membership inference attacks against machine learning models. In Proc. 2017 IEEE Symposium on Security and Privacy (SP) 3&#x2013;18 (IEEE, 2017).\" href=\"#ref-CR2\" id=\"ref-link-section-d97577649e473\">2<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Carlini, N. et al. Membership inference attacks from first principles. In Proc. 2022 IEEE Symposium on Security and Privacy (SP) 1897&#x2013;1914 (IEEE, 2022).\" href=\"#ref-CR3\" id=\"ref-link-section-d97577649e473_1\">3<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Zarifzadeh, S., Liu, P. &amp; Shokri, R. Low-cost high-power membership inference attacks. In Proc. 41st International Conference on Machine Learning 58244&#x2013;58282 (PMLR, 2024).\" href=\"#ref-CR4\" id=\"ref-link-section-d97577649e473_2\">4<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Carlini, N. et al. Extracting training data from large language models. In Proc. 30th USENIX Security Symposium 2633&#x2013;2650 (USENIX, 2021).\" href=\"#ref-CR5\" id=\"ref-link-section-d97577649e473_3\">5<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Carlini, N. et al. Extracting training data from diffusion models. In Proc. 32nd USENIX Security Symposium 5253&#x2013;5270 (USENIX, 2023).\" href=\"#ref-CR6\" id=\"ref-link-section-d97577649e473_4\">6<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 7\" title=\"Nasr, M. et al. Scalable extraction of training data from aligned, production language models. In Proc. Thirteenth International Conference on Learning Representations (ICLR, 2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR7\" id=\"ref-link-section-d97577649e476\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>, which can extract sensitive details about the data used for model training.<\/p>\n<p>Privacy attacks against an AI model can enable detailed inferences about the individuals who contributed to its training data. For example, a membership inference attack (MIA)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 2\" title=\"Shokri, R., Stronati, M., Song, C. &amp; Shmatikov, V. Membership inference attacks against machine learning models. In Proc. 2017 IEEE Symposium on Security and Privacy (SP) 3&#x2013;18 (IEEE, 2017).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR2\" id=\"ref-link-section-d97577649e483\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a> attempts to determine whether the data of a specific patient were included in the training dataset of a model. The extent to which this constitutes a privacy violation is nuanced and depends on factors such as the underlying training population and the deployment context of the model. Although inferring membership for a model trained on a general population may be benign, doing so for a model trained on a narrow, disease- or centre-specific cohort acts as a direct proxy for sensitive medical information. For example, a successful MIA against the model in ref.\u2009<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 10\" title=\"Yoo, S.-K. et al. Prediction of checkpoint inhibitor immunotherapy efficacy for cancer using routine blood tests and clinical data. Nat. Med. 31, 869&#x2013;880 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR10\" id=\"ref-link-section-d97577649e487\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>, which predicts anti-cancer immunotherapy efficacy from routine blood test data, reveals that an individual has cancer.<\/p>\n<p>The accelerating deployment of medical AI models trained on sensitive patient data<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 11\" title=\"FDA. Artificial intelligence-enabled medical devices. &#010;                https:\/\/www.fda.gov\/medical-devices\/software-medical-device-samd\/artificial-intelligence-and-machine-learning-aiml-enabled-medical-devices&#010;                &#010;               (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR11\" id=\"ref-link-section-d97577649e494\" rel=\"nofollow noopener\" target=\"_blank\">11<\/a> calls for rigorous privacy risk assessments. However, previous studies primarily quantified the success rate of MIAs, in aggregate, across all records in a training dataset. This implicitly averages risk across records, thereby obscuring important information on record- and patient-level attack success. Consequently, the risk that an individual faces by contributing their personal data (often multiple records) to an AI training dataset is poorly understood. Given that medical data are a key target for cybercriminals<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 12\" title=\"Seh, A. H. et al. Healthcare data breaches: insights and implications. Healthcare 8, 133 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR12\" id=\"ref-link-section-d97577649e498\" rel=\"nofollow noopener\" target=\"_blank\">12<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 13\" title=\"Albert Haro Abad, S. C. Health Threat Landscape: ENISA Report 2023. Technical report (European Union Agency for Cybersecurity, 2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR13\" id=\"ref-link-section-d97577649e501\" rel=\"nofollow noopener\" target=\"_blank\">13<\/a>, and pseudonymization alone is increasingly recognized as insufficient to prevent the re-identification of individuals in large, high-dimensional datasets<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Narayanan, A. &amp; Shmatikov, V. Robust de-anonymization of large sparse datasets. In Proc. 2008 IEEE Symposium on Security and Privacy (sp 2008) 111&#x2013;125 (IEEE, 2008).\" href=\"#ref-CR14\" id=\"ref-link-section-d97577649e505\">14<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Gadotti, A., Rocher, L., Houssiau, F., Cre&#x163;u, A.-M. &amp; de Montjoye, Y.-A. Anonymization: the imperfect science of using data while preserving privacy. Sci. Adv. 10, eadn7053 (2024).\" href=\"#ref-CR15\" id=\"ref-link-section-d97577649e505_1\">15<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 16\" title=\"Rocher, L., Hendrickx, J. M. &amp; de Montjoye, Y.-A. A scaling law to model the effectiveness of identification techniques. Nat. Commun. 16, 347 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR16\" id=\"ref-link-section-d97577649e508\" rel=\"nofollow noopener\" target=\"_blank\">16<\/a>, there is a need to improve our understanding of the threat that AI privacy attacks pose to individual patients.<\/p>\n<p>Here we show that deploying medical AI models without protective measures can pose substantial privacy risks to individual data-contributing patients. These risks are particularly acute when membership in a training population itself reveals sensitive medical information. Our privacy audit of AI models trained to perform standard diagnostic (supervised classification) tasks quantifies state-of-the-art MIA success<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"Carlini, N. et al. Membership inference attacks from first principles. In Proc. 2022 IEEE Symposium on Security and Privacy (SP) 1897&#x2013;1914 (IEEE, 2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR3\" id=\"ref-link-section-d97577649e515\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Zarifzadeh, S., Liu, P. &amp; Shokri, R. Low-cost high-power membership inference attacks. In Proc. 41st International Conference on Machine Learning 58244&#x2013;58282 (PMLR, 2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR4\" id=\"ref-link-section-d97577649e518\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a> at the resolution of individual data contributors. Using seven large datasets comprising real-world clinical data, including various types of medical images, electrocardiograms and electronic health records, we demonstrate that the success of a\u00a0MIA is unequally distributed among data-contributing patients. We show that this disparity exists at two levels: (1) the individual patient level, at which some patients experience near-perfect attack success, whereas others remain essentially unaffected; and (2) the group level, at which patient groups underrepresented in a training dataset are often overrepresented among records most vulnerable to MIAs.<\/p>\n<p>Together, our results indicate that privacy attacks against AI models may be much more effective at compromising the privacy of individual data contributors than previously thought. This suggests that current AI privacy risk reporting practices may underestimate individual-level risk and thus motivates the integration of mathematically verifiable risk mitigation strategies such as differential privacy (DP) into medical AI model development workflows.<\/p>\n<p>Attacking AI by simple hypothesis tests<\/p>\n<p>A popular deployment strategy for AI models gives users access to a model through a prediction interface, which, for a given input (for example, the chest radiograph of a patient), returns a corresponding prediction (for example, a 78% chance of pneumonia). This black-box access to a model can be exploited by an untrusted user to conduct a MIA that shows the membership status of a target record, that is, whether the target record was a member of the training dataset of a model or not (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1a<\/a>). To infer membership status, MIAs typically make use of the fact that AI models are often slightly more confident about their predictions on training than on non-training data.<\/p>\n<p><b id=\"Fig1\" class=\"c-article-section__figure-caption\" data-test=\"figure-caption-text\">Fig. 1: MIA and evaluation strategies.<\/b><img decoding=\"async\" aria-describedby=\"figure-1-desc\" src=\"https:\/\/www.europesays.com\/uk\/wp-content\/uploads\/2026\/06\/41586_2026_10688_Fig1_HTML.png\" alt=\"Fig. 1: MIA and evaluation strategies.\" loading=\"lazy\" width=\"685\" height=\"544\"\/><\/p>\n<p><b>a<\/b>, Schematic of a MIA, in which an untrusted user, only by observing the predictions of a model, aims to infer whether a specific target record was part of the training dataset. The attack is considered successful if the untrusted user can reliably distinguish between model A and model B, which are identical except for the inclusion and exclusion of the target record in the respective training dataset. <b>b<\/b>,<b>c<\/b>, Attack success can be measured either, in aggregate, across all records in the dataset (<b>b<\/b>) or, more granularly, for each record individually across many target models (<b>c<\/b>).<\/p>\n<p>Likelihood-ratio MIAs\u00a0(LR-MIAs)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"Carlini, N. et al. Membership inference attacks from first principles. In Proc. 2022 IEEE Symposium on Security and Privacy (SP) 1897&#x2013;1914 (IEEE, 2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR3\" id=\"ref-link-section-d97577649e573\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Zarifzadeh, S., Liu, P. &amp; Shokri, R. Low-cost high-power membership inference attacks. In Proc. 41st International Conference on Machine Learning 58244&#x2013;58282 (PMLR, 2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR4\" id=\"ref-link-section-d97577649e576\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>, the current state-of-the-art in MIAs, frame membership inference as a simple vs. simple hypothesis testing problem on the prediction confidence provided by the target model. In essence, LR-MIAs compare the likelihood of the predicted confidence of the target model for the target record under the null (the target record was not a member) and the alternative hypothesis (the target record was a member). Here, the parameters of the distributions under the two hypotheses are specified by parametric fitting of sample confidence values obtained from reference models. Reference models are models assumed to be trained by the attacker and are ideally, but not necessarily, of similar architecture as the target model and trained on data similar to the training dataset of the target model.<\/p>\n<p>Note that objectively larger threats are posed by privacy attacks with stronger assumptions on a potential attacker, such as access to model parameters<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 17\" title=\"Suri, A., Zhang, X., Evans, D. Do parameters reveal more than loss for membership inference? In Proc. 2nd Workshop on High-dimensional Learning Dynamics (HiLD, 2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR17\" id=\"ref-link-section-d97577649e583\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>, access to parameter updates during model training<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 18\" title=\"Geiping, J., Bauermeister, H., Dr&#xF6;ge, H. &amp; Moeller, M. Inverting gradients - how easy is it to break privacy in federated learning? In Proc. 34th International Conference on Neural Information Processing Systems 16937&#x2013;16947 (Curran Associates, 2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR18\" id=\"ref-link-section-d97577649e587\" rel=\"nofollow noopener\" target=\"_blank\">18<\/a> or, furthermore, the ability to modify the model architecture<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 19\" title=\"Fowl, L. H., Geiping, J., Czaja, W., Goldblum, M. &amp; Goldstein, T. Robbing the fed: directly obtaining private data in federated learning with modified models. In Proc. International Conference on Learning Representations (ICLR, 2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR19\" id=\"ref-link-section-d97577649e591\" rel=\"nofollow noopener\" target=\"_blank\">19<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 20\" title=\"Feng, S. &amp; Tram&#xE8;r, F. Privacy backdoors: stealing data with corrupted pretrained models. In Proc. 41st International Conference on Machine Learning 13326&#x2013;13364 (PMLR, 2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR20\" id=\"ref-link-section-d97577649e594\" rel=\"nofollow noopener\" target=\"_blank\">20<\/a>. However, we do not consider them in this study as their strong assumptions are not realistic for careful, practical deployment scenarios. By contrast, the type of attack we consider here requires querying the target model only once (to obtain a prediction for the target record) and may thus be executed by any attacker posing as a real user of an AI system. Notably, as the attacks we study are executed against fully trained models, data-governance-preserving techniques such as federated\/swarm-learning<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 21\" title=\"Warnat-Herresthal, S. et al. Swarm Learning for decentralized and confidential clinical machine learning. Nature 594, 265&#x2013;270 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR21\" id=\"ref-link-section-d97577649e598\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a> provide no protection.<\/p>\n<p>From aggregate to patient-level risk<\/p>\n<p>MIA performance is evaluated through a receiver operating characteristic (ROC) analysis<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"Fawcett, T. An introduction to ROC analysis. Pattern Recognit. Lett. 27, 861&#x2013;874 (2006).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR22\" id=\"ref-link-section-d97577649e610\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a> on numerous repetitions of the MIA game scenario, in which an untrusted user is challenged to guess the membership status of a given record (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1a<\/a>). In practice, owing to the computational cost of training AI models, attack success is typically evaluated using a single target model. More specifically, a target model is trained on a random subset of the training dataset, and subsequently, an ROC analysis is performed on the aggregated membership predictions for all records in the dataset (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1b<\/a>). Although practical, this approach has a key shortcoming: it provides no indication of the performance of the attack for individual records or patients.<\/p>\n<p>To address this issue, we propose a simple technique for estimating record-, and by extension, patient-level vulnerability to LR-MIAs (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1c<\/a>). In brief, using a large set of target models (N\u00a0=\u00a0200) trained on random patient subsets, we estimate, for each training record, sampling distributions of the confidence of the target model under the null and alternative hypotheses in LR-MIAs. In other words, we estimate empirical distributions of confidence values as provided by target models, partitioned into models trained and not trained on the target record. Because these distributions are assumed to take Gaussian form in LR-MIAs<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"Carlini, N. et al. Membership inference attacks from first principles. In Proc. 2022 IEEE Symposium on Security and Privacy (SP) 1897&#x2013;1914 (IEEE, 2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR3\" id=\"ref-link-section-d97577649e629\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Zarifzadeh, S., Liu, P. &amp; Shokri, R. Low-cost high-power membership inference attacks. In Proc. 41st International Conference on Machine Learning 58244&#x2013;58282 (PMLR, 2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR4\" id=\"ref-link-section-d97577649e632\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>, record-level attack success, as measured by the area under the ROC curve (AUC), can be calculated in closed form (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Sec9\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>). A high AUC score, close to the maximum value of 1.0, suggests high privacy risk: a MIA for this record could achieve high sensitivity with little to no false positives. Notably, the record-level MIA AUC also offers a probabilistic interpretation<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"Fawcett, T. An introduction to ROC analysis. Pattern Recognit. Lett. 27, 861&#x2013;874 (2006).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR22\" id=\"ref-link-section-d97577649e639\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a>: the record-level MIA AUC is the probability that a confidence score from a target model trained on the target record is larger than a score from a target model not trained on the target record.<\/p>\n<p>Correctly determining the membership status for one of the records contributed by an individual patient reveals the membership status of the patient. Thus, we compute patient-level scores by taking the maximum across all record-level scores for a given patient. The raw record-level scores and the average patient-level scores can be found in Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>.<\/p>\n<p>Notably, our technique for measuring record-level attack success reduces to estimating the bi-normal AUC from sample statistics and thus has desirable statistical properties. Its standard error at the record level can be computed in closed form<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 23\" title=\"Hanley, J. A. &amp; McNeil, B. J. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology 143, 29&#x2013;36 (1982).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR23\" id=\"ref-link-section-d97577649e652\" rel=\"nofollow noopener\" target=\"_blank\">23<\/a> (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Sec9\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>). As expected, using a total of N\u00a0=\u00a0200 target models evenly split between null and alternative hypotheses for each record, the standard error of the record-level MIA AUC is small across all records in the investigated datasets (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>).<\/p>\n<p>Attacking open-source models<\/p>\n<p>Recent advances in attack design have made LR-MIAs much more practical. To illustrate the practical feasibility of conducting MIAs, we demonstrate attacks against two chest radiograph models from the TorchXrayVision<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 24\" title=\"Cohen, J. P. et al. TorchXRayVision: a library of chest X-ray datasets and models. In Proc. 5th International Conference on Medical Imaging with Deep Learning 231&#x2013;249 (PMLR, 2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR24\" id=\"ref-link-section-d97577649e673\" rel=\"nofollow noopener\" target=\"_blank\">24<\/a> library. We used the Robust Membership Inference Attack<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Zarifzadeh, S., Liu, P. &amp; Shokri, R. Low-cost high-power membership inference attacks. In Proc. 41st International Conference on Machine Learning 58244&#x2013;58282 (PMLR, 2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR4\" id=\"ref-link-section-d97577649e677\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a> (RMIA), an improved LR-MIA that requires only one or two reference models, compared with more than 100 for the Likelihood Ratio Attack<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"Carlini, N. et al. Membership inference attacks from first principles. In Proc. 2022 IEEE Symposium on Security and Privacy (SP) 1897&#x2013;1914 (IEEE, 2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR3\" id=\"ref-link-section-d97577649e681\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a> (LiRA). RMIA achieves this efficiency gain by effectively using reference data (data similar to the target record) alongside the target record to query the target model. Crucially, the attack does not require knowledge of the membership status of the reference data.<\/p>\n<p>We simulated a realistic attack setting in which an attacker lacks access to the training dataset of the target model to train reference models and is further constrained by computational resources. Specifically, we used only a single pre-trained PadChest<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 25\" title=\"Bustos, A., Pertusa, A., Salinas, J.-M. &amp; de la Iglesia-Vay&#xE1;, M. PadChest: a large chest x-ray image dataset with multi-label annotated reports. Med. Image Anal. 66, 101797 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR25\" id=\"ref-link-section-d97577649e688\" rel=\"nofollow noopener\" target=\"_blank\">25<\/a> model as a reference model to perform attacks against the CheXpert<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 26\" title=\"Irvin, J. et al. CheXpert: a large chest radiograph dataset with uncertainty labels and expert comparison. In Proc. Thirty-Third AAAI Conference on Artificial Intelligence 590&#x2013;597 (PKP Publishing, 2019).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR26\" id=\"ref-link-section-d97577649e692\" rel=\"nofollow noopener\" target=\"_blank\">26<\/a> and MIMIC-CXR<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 27\" title=\"Johnson, A. E. W. et al. MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Sci. Data 6, 317 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR27\" id=\"ref-link-section-d97577649e696\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a> models of the library. In this setting, also known as an offline attack, an attacker incurs no computational cost in training the reference model. Instead, they simply need to obtain predictions from the reference model for both the target record and the reference data. This can be done efficiently on commodity hardware without a graphics processing unit (GPU). To conduct the attack, we queried the target model once to collect confidence values for all target records. Using this collection, we then computed RMIA test statistics for each target record by randomly selecting, independent of membership status, N\u00a0=\u00a0500 confidence values from the other targets in this collection as reference data. This strategy would effectively conceal the additional reference data queries to the target model in a real attack.<\/p>\n<p>We evaluated attack success on a combined dataset of records from CheXpert and MIMIC-CXR (N\u00a0=\u00a025,\u00a0000 each), which were, respectively, labelled as members and non-members for the CheXpert model (v.v. for the MIMIC-CXR model). In this setting, RMIA achieved substantial aggregate success with respective AUC scores of 0.61 and 0.65 (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a>). Note that owing to the distribution shift between members and non-members, results from this evaluation setting are not directly comparable to the standard evaluation protocol in which members and non-members are sampled at random from the training dataset. Notably, however, such a distribution shift is expected in a real attack, and this setting is thus of high interest.<\/p>\n<p><b id=\"Fig2\" class=\"c-article-section__figure-caption\" data-test=\"figure-caption-text\">Fig. 2: MIAs pose substantial privacy risks to individual data-contributing patients.<\/b><img decoding=\"async\" aria-describedby=\"figure-2-desc\" src=\"https:\/\/www.europesays.com\/uk\/wp-content\/uploads\/2026\/06\/41586_2026_10688_Fig2_HTML.png\" alt=\"Fig. 2: MIAs pose substantial privacy risks to individual data-contributing patients.\" loading=\"lazy\" width=\"685\" height=\"667\"\/><\/p>\n<p><b>a<\/b>, Aggregate MIA success for a realistic attack against open-source CheXpert and MIMIC-CXR models (RMIA offline, R\u00a0=\u00a01 reference model pre-trained on PadChest). <b>b<\/b>, eSF analysis of patient-level MIA AUC scores computed using N\u00a0=\u00a0200 target models for each dataset (residual networks, about 1.5 million parameters each). Patient-level scores are computed as the maximum record-level score for a given patient. <b>c<\/b>, ROC analysis of aggregate attack success (LiRA online, vertical-average mean for N\u00a0=\u00a010 target models, R\u00a0=\u00a0190 reference models). <b>d<\/b>,<b>e<\/b>, eSF plots of patient-level MIA AUC scores alongside diagnostic performance (macro-average AUC) on unseen test data for N\u00a0=\u00a0200 target models with varying levels of record-level (\u03b5,\u00a0\u03b4)-DP privacy protection: PTB-XL (<b>d<\/b>) and EMBED (<b>e<\/b>). \u03b4 was kept constant at 1\/D, where D is the dataset size. <b>f<\/b>,<b>g<\/b>, eSF plots of patient-level MIA AUC scores alongside diagnostic performance (macro-average AUC) on unseen test data for N\u00a0=\u00a0200 target models of increasing model capacity: Fitzpatrick-17k (<b>f<\/b>) and CheXpert (<b>g<\/b>). Error bars indicate s.d., round brackets indicate aggregate attack AUC; square brackets indicate AUC upper bound implied by record-level DP accounting; dashed grey lines in ROC curve plots indicate random-guessing performance; and dashed lines in eSF plots indicate 95% Greenwood CI.<\/p>\n<p>Near-perfect success for some patients<\/p>\n<p>After demonstrating realistic attacks against two open-source models, we next investigated how effectively MIAs can compromise the privacy of individual patients. To this end, we measured patient-level MIA success across a diverse range of medical datasets using, for each, a large set of target models. Notably, we used state-of-the-art model training techniques (for example, data augmentation, weight decay\u00a0and learning rate schedules) and furthermore, took explicit countermeasures to prevent overfitting, which is known to exacerbate privacy risks<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 2\" title=\"Shokri, R., Stronati, M., Song, C. &amp; Shmatikov, V. Membership inference attacks against machine learning models. In Proc. 2017 IEEE Symposium on Security and Privacy (SP) 3&#x2013;18 (IEEE, 2017).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR2\" id=\"ref-link-section-d97577649e808\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 28\" title=\"Yeom, S., Giacomelli, I., Fredrikson, M. &amp; Jha, S. Privacy risk in machine learning: analyzing the connection to overfitting. In Proc. 2018 IEEE 31st Computer Security Foundations Symposium (CSF) 268&#x2013;282 (IEEE, 2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR28\" id=\"ref-link-section-d97577649e811\" rel=\"nofollow noopener\" target=\"_blank\">28<\/a>. As a result, the investigated target models, despite being trained on roughly half of the available data each, provide high diagnostic performance within a few percentage points of published baselines (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Sec9\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>).<\/p>\n<p>Across all investigated datasets and models, we identified a small subset of patients who are highly vulnerable to LR-MIAs. This is indicated by empirical survival functions (eSF) of patient-level MIA AUC scores, which, for a given score, show the proportion of patients with this score or higher (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2b<\/a>). By contrast, ROC curves of aggregate attack success and their corresponding AUC scores do not deviate substantially from the random-guessing baseline, thus incorrectly indicating a low attack vulnerability (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2c<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">1c\u2013e<\/a>). This suggests that average-case metrics of attack success, as used in the standard evaluation protocol, are unsuitable measures of privacy risk. They do not accurately reflect that some records or patients may be highly vulnerable, whereas the vast majority are not.<\/p>\n<p>For the two non-imaging datasets, MIMIC-IV-ED<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 29\" title=\"Johnson, A. et al. MIMIC-IV-ED (v.1.0) (PhysioNet, 2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR29\" id=\"ref-link-section-d97577649e833\" rel=\"nofollow noopener\" target=\"_blank\">29<\/a> (electronic health records) and PTB-XL<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 30\" title=\"Wagner, P. et al. PTB-XL, a large publicly available electrocardiography dataset. Sci. Data 7, 154 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR30\" id=\"ref-link-section-d97577649e837\" rel=\"nofollow noopener\" target=\"_blank\">30<\/a> (electrocardiograms), we simulated attack settings in which an attacker only has partial access to the target record (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig6\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>). Although MIA success generally decreases under partial data access, a subset of patients retain high AUC scores, even in settings in which the attacker has access to only basic clinical information\u2014such as a patients\u2019 age, sex, chief complaints and vital signs (MIMIC-IV-ED), or only the lead I signal from a 12-lead electrocardiogram (PTB-XL).<\/p>\n<p>We verified how resolvable the discovered vulnerabilities are by training models with different levels of record-level (\u03b5,\u00a0\u03b4)-DP protection (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2d,e<\/a>). As expected, we find that patient-level MIA risk decreases with stronger levels of privacy protection (smaller \u03b5 values). Moreover, in most scenarios, we observe no violation of the record-level DP guarantee (indicated by the square brackets in the panel legend), although many patients contributed multiple records. Violations are observed only for a subset of patients under strong privacy protection (\u03b5\u00a0=\u00a01), in which some patients have MIA AUC scores exceeding the upper bound on the MIA AUC implied by the record-level DP guarantee. This behaviour is expected and could be mitigated by implementing patient-level DP accounting.<\/p>\n<p>Larger models, greater risks<\/p>\n<p>Many of the recent AI success stories have been driven not by methodological advances but by scaling up model and dataset sizes<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 31\" title=\"Kaplan, J. et al. Scaling laws for neural language models. Preprint at &#010;                https:\/\/doi.org\/10.48550\/arXiv.2001.08361&#010;                &#010;               (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR31\" id=\"ref-link-section-d97577649e870\" rel=\"nofollow noopener\" target=\"_blank\">31<\/a>. In light of this scaling trend, we next investigated the impact of model capacity on MIA success. For Fitzpatrick 17k<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 32\" title=\"Groh, M. et al. Evaluating deep neural networks trained on clinical images in dermatology with the fitzpatrick 17k dataset. In Proc. 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) 1820&#x2013;1828 (CVPRW, 2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR32\" id=\"ref-link-section-d97577649e874\" rel=\"nofollow noopener\" target=\"_blank\">32<\/a> and CheXpert, we trained models with increasing capacity, including wide residual networks<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 33\" title=\"Zagoruyko, S. &amp; Komodakis, N. Wide residual networks. In Proc. British Machine Vision Conference (BMVC) (eds Wilson, R. C. et al.) 87.1&#x2013;87.12 (BMVA Press, 2016).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR33\" id=\"ref-link-section-d97577649e878\" rel=\"nofollow noopener\" target=\"_blank\">33<\/a> (WRN-28-2 and WRN-40-4) and vision transformers<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 34\" title=\"Dosovitskiy, A. et al. An image is worth 16x16 words: transformers for image recognition at scale. In Proc. IEEE\/CVF Conference on Computer Vision and Pattern Recognition 45&#x2013;67 (ICLR, 2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR34\" id=\"ref-link-section-d97577649e882\" rel=\"nofollow noopener\" target=\"_blank\">34<\/a> (ViT-B\/16 and ViT-L\/16). Where computationally feasible, vision transformers were trained on images of different sizes: 64\u00a0\u00d7\u00a064 and 128\u00a0\u00d7\u00a0128 pixels; this is indicated by a trailing number behind the model name (for example, ViT-B\/16-64 and ViT-B\/16-128).<\/p>\n<p>We find that MIA success (both at the aggregate and patient levels) increases with model capacity. We observe that the relative share of patients highly vulnerable to MIAs increases greatly for larger models, often by an order of magnitude (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2f,g<\/a>). For the dermatology dataset (Fitzpatrick 17k), increasing model capacity yields large gains in diagnostic performance with a pronounced increase between WRN-40-4 and ViT-B\/16-128, which was pre-trained on a large dataset of more than 14 million natural images. However, simultaneously, the number of patients with near-perfect attack success (AUC score of 0.95 or higher) increases substantially: 0 (WRN-28-2), 1 out of 10,000 (WRN-40-4), 1 out of 1,000 (ViT-B\/16-64) and 1 out of 10 (ViT-B\/16-128). We observe a similar trend in the much larger dataset CheXpert, although attack success is generally lower. Notably, for CheXpert, vision transformer models do not achieve diagnostic performance competitive with WRN-based models. This is probably because of the diminished utility of natural-image pre-training for medical greyscale images<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 35\" title=\"Raghu, M., Zhang, C., Kleinberg, J. &amp; Bengio, S. Transfusion: understanding transfer learning for medical imaging. In Proc. 33rd International Conference on Neural Information Processing Systems 3347&#x2013;3357 (NIPS, 2019).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR35\" id=\"ref-link-section-d97577649e892\" rel=\"nofollow noopener\" target=\"_blank\">35<\/a>.<\/p>\n<p>Attack success varies by subgroup<\/p>\n<p>Motivated by recent findings<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 36\" title=\"Seyyed-Kalantari, L., Zhang, H., McDermott, M. B. A., Chen, I. Y. &amp; Ghassemi, M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat. Med. 27, 2176&#x2013;2182 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR36\" id=\"ref-link-section-d97577649e905\" rel=\"nofollow noopener\" target=\"_blank\">36<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 37\" title=\"Daneshjou, R. et al. Disparities in dermatology AI performance on a diverse, curated clinical image set. Sci. Adv. 8, eabq6147 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR37\" id=\"ref-link-section-d97577649e908\" rel=\"nofollow noopener\" target=\"_blank\">37<\/a>, which revealed that the diagnostic performance of AI models can differ across patient subgroups, we investigated whether differences in privacy risk exist between subgroups. To this end, we focused our analysis on the most vulnerable records (99th MIA AUC percentile) and compared how frequently a subgroup appears in this extreme-risk tail compared with the overall dataset. We did not consider differences in aggregate attack success, as we previously identified this metric as an unsuitable measure of privacy risk.<\/p>\n<p>We find that extreme MIA risk is unequally distributed across patient subgroups when stratifying by disease status, self-reported race, sex, imaging protocol or health insurance. More precisely, for most comparisons, we observe significant differences in subgroup composition between the most vulnerable records and the overall dataset (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>). For example, in MIMIC-IV-ED, records from Black patients, patients with Medicaid insurance or patients diagnosed with cancer were observed more frequently than expected among the most vulnerable records (+31%, +126%, and +18% relative change to the overall dataset, respectively). Raw data on the composition of the extreme MIA risk tails as well as the overall datasets are provided in Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>\u2013<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">16<\/a>. To find factors that could explain the observed differences, we performed a post hoc test analysis and computed Pearson residuals for all subgroup comparisons (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>).<\/p>\n<p><b id=\"Fig3\" class=\"c-article-section__figure-caption\" data-test=\"figure-caption-text\">Fig. 3: Significant differences in extreme MIA risk between patient subgroups.<\/b><img decoding=\"async\" aria-describedby=\"figure-3-desc\" src=\"https:\/\/www.europesays.com\/uk\/wp-content\/uploads\/2026\/06\/41586_2026_10688_Fig3_HTML.png\" alt=\"Fig. 3: Significant differences in extreme MIA risk between patient subgroups.\" loading=\"lazy\" width=\"685\" height=\"727\"\/><\/p>\n<p><b>a\u2013d<\/b>, The panel rows show two-sided \u03c72-test results for subgroup counts among the 99th record-level MIA AUC percentile for CheXpert (<b>a<\/b>), MIMIC-CXR (<b>b<\/b>), EMBED (<b>c<\/b>) and MIMIC-IV-ED (<b>d<\/b>). Pearson residuals measure the contribution of a group to the test statistic. A large, positive value indicates that more records were observed for this group than expected. A negative value indicates the opposite. The column titles show stratification variables; bars are coloured according to the relative share of records of a group in the training dataset. *P\u2009\u2264\u20090.05, **P\u2009\u2264\u20090.01, ***P\u2009\u2264\u20090.001 and NS, not significant. Multiple comparison correction was applied row-wise using the Bonferroni method; statistical significance could not be tested for CheXpert and MIMIC-CXR disease label groups as the categorization is not mutually exclusive; race subgroup comparisons in EMBED are likely confounded by breast density differences. Left to right, CheXpert\/MIMIC-CXR disease labels refer to no finding\u00a0(NF), enlarged cardiomediastinum\u00a0(EC), cardiomegaly\u00a0(Cm), lung opacity\u00a0(LO), lung lesion\u00a0(LL), oedema\u00a0(Ed), consolidation\u00a0(Co), pneumonia\u00a0(Pn), atelectasis\u00a0(At), pneumothorax\u00a0(Px), pleural effusion\u00a0(PE), pleural other\u00a0(PO), fracture\u00a0(Fr) and support devices\u00a0(SD). Imaging protocol abbreviations refer to anteroposterior\u00a0(AP), posteroanterior\u00a0(PA), lateral\u00a0(L), lateral left\u00a0(LL), mediolateral oblique\u00a0(MLO) and craniocaudal\u00a0(CC). BI-RADS indicates breast cancer assessment: incomplete (BI-RADS-0) to biopsy-proven malignancy (BI-RADS-6) and breast density: mostly fatty (BI-RADS-A) to extremely dense (BI-RADS-D). Left to right, adjusted P values are as follows. CheXpert: 8.1\u00a0\u00d7\u00a010\u22129, 9.7\u00a0\u00d7\u00a010\u22122, 2.1\u00a0\u00d7\u00a010\u22125; MIMIC-CXR: 0.1, 2.2\u00a0\u00d7\u00a010\u221223, 1.6\u00a0\u00d7\u00a010\u221229; EMBED: 1.0\u00a0\u00d7\u00a010\u221268, &lt;1.0\u00a0\u00d7\u00a010\u2212100, 1.0\u00a0\u00d7\u00a010\u22128, 1.0, 2.7\u00a0\u00d7\u00a010\u22127. MIMIC-IV-ED: 0.61, 4.1\u00a0\u00d7\u00a010\u221224, 2.3\u00a0\u00d7\u00a010\u22129, 6.3\u00a0\u00d7\u00a010\u221220, 1.8\u00a0\u00d7\u00a010\u22122, 2.3\u00a0\u00d7\u00a010\u221253.<\/p>\n<p>We primarily observe large, positive Pearson residuals for underrepresented groups in the datasets, suggesting that relative group size influences MIA risk. Consider, for example, EMBED<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 38\" title=\"Jeong, J. J. et al. The EMory BrEast imaging Dataset (EMBED): a racially diverse, granular dataset of 3.4 million screening and diagnostic mammographic images. Radiol. Artif. Intell. 5, e220047 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR38\" id=\"ref-link-section-d97577649e1019\" rel=\"nofollow noopener\" target=\"_blank\">38<\/a>, a mammography dataset comprising mostly negative findings, that is, unremarkable mammograms of healthy breasts with no indication of a tumour. Models for this dataset are trained to predict breast density, and thus never have direct access to tumour findings. Despite this, benign tumour findings (BI-RADS-2) and tumour findings suspicious of malignancy (BI-RADS-4) account for a disproportionately large share of the most vulnerable records (+60% and +1,179% relative change to the overall dataset, respectively). Similarly, otherwise relatively uncommon images of almost entirely fatty (BI-RADS-A) or extremely dense (BI-RADS-D) breasts also occur disproportionately frequently (+90% and +755%, respectively).<\/p>\n<p>To further investigate the relationship between group size and MIA risk, we conducted a meta-analysis of all computed Pearson residuals (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#Fig8\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a>). Confirming previous observations, we find that large positive Pearson residuals occur mostly for small groups (those that contribute less than 20% of the records of a dataset). Moreover, we observe a weak to moderate negative correlation between group size and Pearson residuals. This suggests that the observed differences in MIA risk may, at least in part, be driven by group-size differences in the training data.<\/p>\n<p>Discussion<\/p>\n<p>We present data from the first patient-level privacy audit of medical AI models. Our findings confirm early observations of MIA risk heterogeneity<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Long, Y. et al. A pragmatic approach to membership inferences on machine learning models. In Proc. 2020 IEEE European Symposium on Security and Privacy (EuroS&amp;P) 521&#x2013;534 (IEEE, 2020).\" href=\"#ref-CR39\" id=\"ref-link-section-d97577649e1037\">39<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Aerni, M., Zhang, J. &amp; Tram&#xE8;r, F. Evaluations of machine learning privacy defenses are misleading. In Proc. 2024 on ACM SIGSAC Conference on Computer and Communications Security 1271&#x2013;1284 (CCS, 2024).\" href=\"#ref-CR40\" id=\"ref-link-section-d97577649e1037_1\">40<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Kulynych, B., Yaghini, M. &amp; Cherubin, G., Veale, M., Troncoso, C. Disparate vulnerability to membership inference attacks. In Proc. Privacy Enhancing Technologies 460&#x2013;480 (sciendo, 2022).\" href=\"#ref-CR41\" id=\"ref-link-section-d97577649e1037_2\">41<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 42\" title=\"Chang, H. &amp; Shokri, R. On the privacy risks of algorithmic fairness. In Proc. 2021 IEEE European Symposium on Security and Privacy (EuroS&amp;P) 292&#x2013;303 (IEEE, 2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR42\" id=\"ref-link-section-d97577649e1040\" rel=\"nofollow noopener\" target=\"_blank\">42<\/a> and, at the same time, substantially advance previous AI privacy auditing efforts along three key dimensions. First, our work marks a shift towards patient-level risk assessment, which is crucial for real-world clinical datasets, in which individuals often contribute multiple, similar records. Second, we demonstrate that aggregate success rates, as used in the standard evaluation protocol and previous subgroup analyses<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 41\" title=\"Kulynych, B., Yaghini, M. &amp; Cherubin, G., Veale, M., Troncoso, C. Disparate vulnerability to membership inference attacks. In Proc. Privacy Enhancing Technologies 460&#x2013;480 (sciendo, 2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR41\" id=\"ref-link-section-d97577649e1044\" rel=\"nofollow noopener\" target=\"_blank\">41<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 42\" title=\"Chang, H. &amp; Shokri, R. On the privacy risks of algorithmic fairness. In Proc. 2021 IEEE European Symposium on Security and Privacy (EuroS&amp;P) 292&#x2013;303 (IEEE, 2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR42\" id=\"ref-link-section-d97577649e1047\" rel=\"nofollow noopener\" target=\"_blank\">42<\/a>, underestimate true privacy risks. Third, we confirm that MIA vulnerabilities previously observed on low-dimensional benchmark datasets<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Shokri, R., Stronati, M., Song, C. &amp; Shmatikov, V. Membership inference attacks against machine learning models. In Proc. 2017 IEEE Symposium on Security and Privacy (SP) 3&#x2013;18 (IEEE, 2017).\" href=\"#ref-CR2\" id=\"ref-link-section-d97577649e1051\">2<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Carlini, N. et al. Membership inference attacks from first principles. In Proc. 2022 IEEE Symposium on Security and Privacy (SP) 1897&#x2013;1914 (IEEE, 2022).\" href=\"#ref-CR3\" id=\"ref-link-section-d97577649e1051_1\">3<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Zarifzadeh, S., Liu, P. &amp; Shokri, R. Low-cost high-power membership inference attacks. In Proc. 41st International Conference on Machine Learning 58244&#x2013;58282 (PMLR, 2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR4\" id=\"ref-link-section-d97577649e1054\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Long, Y. et al. A pragmatic approach to membership inferences on machine learning models. In Proc. 2020 IEEE European Symposium on Security and Privacy (EuroS&amp;P) 521&#x2013;534 (IEEE, 2020).\" href=\"#ref-CR39\" id=\"ref-link-section-d97577649e1057\">39<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Aerni, M., Zhang, J. &amp; Tram&#xE8;r, F. Evaluations of machine learning privacy defenses are misleading. In Proc. 2024 on ACM SIGSAC Conference on Computer and Communications Security 1271&#x2013;1284 (CCS, 2024).\" href=\"#ref-CR40\" id=\"ref-link-section-d97577649e1057_1\">40<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Kulynych, B., Yaghini, M. &amp; Cherubin, G., Veale, M., Troncoso, C. Disparate vulnerability to membership inference attacks. In Proc. Privacy Enhancing Technologies 460&#x2013;480 (sciendo, 2022).\" href=\"#ref-CR41\" id=\"ref-link-section-d97577649e1057_2\">41<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 42\" title=\"Chang, H. &amp; Shokri, R. On the privacy risks of algorithmic fairness. In Proc. 2021 IEEE European Symposium on Security and Privacy (EuroS&amp;P) 292&#x2013;303 (IEEE, 2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR42\" id=\"ref-link-section-d97577649e1060\" rel=\"nofollow noopener\" target=\"_blank\">42<\/a> are present, and arguably more critical, in large representative clinical datasets. Below, we briefly discuss our findings and their implications.<\/p>\n<p>The fact that MIAs can achieve near-perfect success rates for individual patients is not adequately captured by the standard evaluation protocol, which measures attack success in aggregate across records. This remains true even when evaluating aggregate attack success at very low false-positive rates (for example, 10\u22124), which is the current standard practice. Thus, reporting standards for AI privacy audits need to change. Audits should report the success of privacy attacks at the level of individual data contributors or, if the necessary patient- or person-level identifiers are unavailable, at the record level.<\/p>\n<p>We observed that the number of patients highly vulnerable to MIAs increases drastically for larger models. Although the magnitude of this change in patient-level risk was previously unknown, other works have also reported greater attack success against larger, more performant models<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"Carlini, N. et al. Membership inference attacks from first principles. In Proc. 2022 IEEE Symposium on Security and Privacy (SP) 1897&#x2013;1914 (IEEE, 2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR3\" id=\"ref-link-section-d97577649e1072\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 5\" title=\"Carlini, N. et al. Extracting training data from large language models. In Proc. 30th USENIX Security Symposium 2633&#x2013;2650 (USENIX, 2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR5\" id=\"ref-link-section-d97577649e1075\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 7\" title=\"Nasr, M. et al. Scalable extraction of training data from aligned, production language models. In Proc. Thirteenth International Conference on Learning Representations (ICLR, 2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR7\" id=\"ref-link-section-d97577649e1078\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 43\" title=\"Carlini, N. et al. Quantifying memorization across neural language models. In Proc. Eleventh International Conference on Learning Representations (ICLR, 2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR43\" id=\"ref-link-section-d97577649e1081\" rel=\"nofollow noopener\" target=\"_blank\">43<\/a>. This observation that privacy risks grow with model size and predictive performance is explained by theoretical research<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 44\" title=\"Feldman, V. Does learning require memorization? a short tale about a long tail. In Proc. 52nd Annual ACM SIGACT Symposium on Theory of Computing 954&#x2013;959 (STOC, 2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR44\" id=\"ref-link-section-d97577649e1085\" rel=\"nofollow noopener\" target=\"_blank\">44<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 45\" title=\"Feldman, V. &amp; Zhang, C. What neural networks memorize and why: discovering the long tail via influence estimation. In Proc. 34th International Conference on Neural Information Processing Systems 2881&#x2013;2891 (NIPS, 2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR45\" id=\"ref-link-section-d97577649e1088\" rel=\"nofollow noopener\" target=\"_blank\">45<\/a>, which postulates that, for long-tailed data distributions, fitting atypical records from the tail is necessary to achieve optimal performance on unseen data at test time. Our results provide further empirical support for this theory and, together, suggest that a trade-off between patient privacy and model performance is inevitable, particularly for rare diseases. Generally, as we found that the number of patients highly vulnerable to MIAs increases by orders of magnitude with larger models, we recommend carefully evaluating the need for the performance improvements they offer.<\/p>\n<p>We found substantial differences in the frequency with which patients from different subgroups experience extreme MIA risk. The fact that some of these groups (for example, self-reported race subgroups in chest radiographs) are not readily distinguishable by human experts raises concerns that MIA risk differences, which probably exist beyond the stratification variables we investigated, may pass unnoticed in practice. We found that the observed risk differences are driven, at least in part, by group-size differences in the training data. Groups of patients that are underrepresented in a model training dataset are often overrepresented among the records most susceptible to MIAs. By contrast, the opposite often holds for majority groups. This finding\u2014that a disproportionately large share of the AI privacy risk burden rests on underrepresented groups\u2014complements the existing literature on health inequalities, which has reported worse health outcomes and life expectancy for marginalized and minority groups<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 46\" title=\" World Health Organization et al. World Report on Social Determinants of Health Equity, 2025 (WHO, 2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR46\" id=\"ref-link-section-d97577649e1095\" rel=\"nofollow noopener\" target=\"_blank\">46<\/a>. Our findings suggest that current trends in medical AI development and deployment could exacerbate these health inequalities. Previous research has shown that the diagnostic performance of AI models, which typically increases with the amount of suitable training data, can be significantly lower for underrepresented (minority) groups<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 36\" title=\"Seyyed-Kalantari, L., Zhang, H., McDermott, M. B. A., Chen, I. Y. &amp; Ghassemi, M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat. Med. 27, 2176&#x2013;2182 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR36\" id=\"ref-link-section-d97577649e1099\" rel=\"nofollow noopener\" target=\"_blank\">36<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 37\" title=\"Daneshjou, R. et al. Disparities in dermatology AI performance on a diverse, curated clinical image set. Sci. Adv. 8, eabq6147 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR37\" id=\"ref-link-section-d97577649e1102\" rel=\"nofollow noopener\" target=\"_blank\">37<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 47\" title=\"Obermeyer, Z., Powers, B., Vogeli, C. &amp; Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations. Science 366, 447&#x2013;453 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR47\" id=\"ref-link-section-d97577649e1105\" rel=\"nofollow noopener\" target=\"_blank\">47<\/a>. Thus, there is a possibility of a vicious cycle in which minority groups place decreasing levels of trust in AI model performance and security, leading to a decreased willingness to contribute to model training datasets.<\/p>\n<p>MIAs facilitate data extraction attacks against generative AI models<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Carlini, N. et al. Extracting training data from large language models. In Proc. 30th USENIX Security Symposium 2633&#x2013;2650 (USENIX, 2021).\" href=\"#ref-CR5\" id=\"ref-link-section-d97577649e1113\">5<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Carlini, N. et al. Extracting training data from diffusion models. In Proc. 32nd USENIX Security Symposium 5253&#x2013;5270 (USENIX, 2023).\" href=\"#ref-CR6\" id=\"ref-link-section-d97577649e1113_1\">6<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 7\" title=\"Nasr, M. et al. Scalable extraction of training data from aligned, production language models. In Proc. Thirteenth International Conference on Learning Representations (ICLR, 2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR7\" id=\"ref-link-section-d97577649e1116\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>. Thus, our findings have potentially far-reaching implications for generative AI privacy risk assessments. Extraction attacks allow for high-fidelity reconstruction of full individual records from the training dataset of a model and have been demonstrated for large language models<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 5\" title=\"Carlini, N. et al. Extracting training data from large language models. In Proc. 30th USENIX Security Symposium 2633&#x2013;2650 (USENIX, 2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR5\" id=\"ref-link-section-d97577649e1120\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a>, diffusion-based image generation models<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 6\" title=\"Carlini, N. et al. Extracting training data from diffusion models. In Proc. 32nd USENIX Security Symposium 5253&#x2013;5270 (USENIX, 2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR6\" id=\"ref-link-section-d97577649e1124\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a> and recently, aligned, production-level large language models<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 7\" title=\"Nasr, M. et al. Scalable extraction of training data from aligned, production language models. In Proc. Thirteenth International Conference on Learning Representations (ICLR, 2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR7\" id=\"ref-link-section-d97577649e1128\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>. Although our study focused on discriminative (diagnostic) AI models, the type of attack we studied is generally applicable and can be used against generative models with little to no modification. We thus see the exploration of our proposed methodology for estimating record- and patient-level MIA success against generative models as an interesting direction for future research. Given the substantial computational resources this would require, exploring scalable approximation techniques is another valuable avenue to investigate.<\/p>\n<p>Unlocking the full potential of medical AI will require training models on vast medical datasets; this depends on gaining and upholding the trust of data-contributing patients. To this end, mathematically verifiable approaches to risk mitigation, such as DP<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 48\" title=\"Dwork, C. &amp; Roth, A. The algorithmic foundations of differential privacy. Foundations Trends Theoret. Comput. Sci. 9, 211&#x2013;487 (2014).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR48\" id=\"ref-link-section-d97577649e1135\" rel=\"nofollow noopener\" target=\"_blank\">48<\/a>, are emerging as the most promising solution. DP, by carefully perturbing parameter updates with white noise during model training or fine-tuning<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 49\" title=\"Abadi, M. et al. Deep learning with differential privacy. In Proc. 2016 ACM SIGSAC Conference on Computer and Communications Security 308&#x2013;318 (CCS, 2016).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR49\" id=\"ref-link-section-d97577649e1139\" rel=\"nofollow noopener\" target=\"_blank\">49<\/a>, limits the contribution of the data of any individual to the parameter update and, by extension, to the final model. This provably protects the privacy of any data-contributing patient, no matter how unique or atypical their data may be. Our experimental data confirmed that stronger levels of DP protection effectively reduce MIA success for all data-contributing patients. However, we also observed that mitigating MIAs requires stronger levels of DP protection than previously thought<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"Carlini, N. et al. Membership inference attacks from first principles. In Proc. 2022 IEEE Symposium on Security and Privacy (SP) 1897&#x2013;1914 (IEEE, 2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR3\" id=\"ref-link-section-d97577649e1143\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 50\" title=\"Nasr, M., Songi, S., Thakurta, A., Papernot, N. &amp; Carlini, N. Adversary instantiation: lower bounds for differentially private machine learning. In Proc. 2021 IEEE Symposium on Security and Privacy (SP) 866&#x2013;882 (IEEE, 2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR50\" id=\"ref-link-section-d97577649e1146\" rel=\"nofollow noopener\" target=\"_blank\">50<\/a>. Specifically, our results indicate that fully mitigating MIAs for all data-contributing patients requires implementing DP protection at the patient level rather than at the record level. Recent research has demonstrated that, in practice, AI models can be trained with strong privacy guarantees while incurring minimal degradation in predictive performance compared with a non-private model<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Ziller, A. et al. Reconciling privacy and accuracy in AI for medical imaging. Nat. Mach. Intell. 6, 764&#x2013;774 (2024).\" href=\"#ref-CR51\" id=\"ref-link-section-d97577649e1150\">51<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Berrada, L. et al. Unlocking accuracy and fairness in differentially private image classification. Preprint at &#10;                https:\/\/doi.org\/10.48550\/arXiv.2308.10888&#10;                &#10;               (2023).\" href=\"#ref-CR52\" id=\"ref-link-section-d97577649e1150_1\">52<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"De, S., Berrada, L., Hayes, J., Smith, S. L. &amp; Balle, B. Unlocking high-accuracy differentially private image classification through scale. Preprint at &#10;                https:\/\/doi.org\/10.48550\/arXiv.2204.13650&#10;                &#10;               (2022).\" href=\"#ref-CR53\" id=\"ref-link-section-d97577649e1150_2\">53<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 54\" title=\"Mckenna, R. et al. Scaling laws for differentially private language models. In Proc. 42nd International Conference on Machine Learning 43375&#x2013;43398 (PMLR, 2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10688-0#ref-CR54\" id=\"ref-link-section-d97577649e1153\" rel=\"nofollow noopener\" target=\"_blank\">54<\/a>. We are thus optimistic that medical AI models protected by DP will have a significant positive impact on health outcomes globally without endangering the privacy of any data-contributing patient.<\/p>\n<p>In summary, we present evidence that MIAs can be highly effective at compromising the privacy of individual data-contributing patients. Given this vulnerability, medical AI models and their deployment contexts should be assessed for the sensitive information that attackers could obtain by successfully inferring training dataset membership. To prevent privacy harm, we recommend that vulnerable models be protected by verifiable risk mitigation strategies and\/or strict access controls.<\/p>\n","protected":false},"excerpt":{"rendered":"Medical artificial intelligence (AI) has immense potential to improve health outcomes, particularly in regions in which specialized medical&hellip;\n","protected":false},"author":2,"featured_media":1048284,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","_share_on_mastodon":"0"},"categories":[6],"tags":[51,18852,3965,477,20468,3966,2753,70,16,15],"class_list":["post-1048283","post","type-post","status-publish","format-standard","has-post-thumbnail","category-business","tag-business","tag-computational-science","tag-humanities-and-social-sciences","tag-information-technology","tag-medical-imaging","tag-multidisciplinary","tag-risk-factors","tag-science","tag-uk","tag-united-kingdom"],"share_on_mastodon":{"url":"","error":""},"_links":{"self":[{"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/posts\/1048283","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/comments?post=1048283"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/posts\/1048283\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/media\/1048284"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/media?parent=1048283"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/categories?post=1048283"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/uk\/wp-json\/wp\/v2\/tags?post=1048283"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}