Menachemi, N. & Collum, T. H. Benefits and drawbacks of electronic health record systems. Risk Manag. Healthc. Policy 4, 47–55 (2011).
Jha, A. K. et al. Use of electronic health records in U.S. hospitals. N. Engl. J. Med. 360, 1628–1638 (2009).
Pendergrass, S. A. & Crawford, D. C. Using electronic health records to generate phenotypes for research. Curr. Protoc. Hum. Genet. 100, e80 (2019).
Chishtie, J. et al. Use of epic electronic health record system for health care research: scoping review. J. Med. Internet Res. 25, e51003 (2023).
Tam, V. et al. Benefits and limitations of genome-wide association studies. Nat. Rev. Genet. 20, 467–484 (2019).
Uffelmann, E. et al. Genome-wide association studies. Nat. Rev. Methods Primers 1, 59 (2021).
Verma, A. et al. PheWAS and Beyond: the landscape of associations with medical diagnoses and clinical measures across 38,662 individuals from Geisinger. Am. J. Hum. Genet. 102, 592–608 (2018).
Sun, J. et al. Translating polygenic risk scores for clinical use by estimating the confidence bounds of risk prediction. Nat. Commun. 12, 5276 (2021).
Dudbridge, F. Power and predictive accuracy of polygenic risk scores. PLoS Genet. 9, e1003348 (2013).
Jayasinghe, D., Eshetie, S., Beckmann, K., Benyamin, B. & Lee, S. H. Advancements and limitations in polygenic risk score methods for genomic prediction: a scoping review. Hum. Genet. 143, 1401–1431 (2024).
Kirchler, M. et al. Large language models improve transferability of electronic health record-based predictions across countries and coding systems. npj Digit. Med. 9, 177 (2026).
Meng, X. et al. The application of large language models in medicine: a scoping review. iScience 27, 109713 (2024).
Liao, K. P. et al. Development of phenotype algorithms using electronic medical records and incorporating natural language processing. BMJ 350, h1885 (2015).
Guevara, M. et al. Large language models to identify social determinants of health in electronic health records. npj Digit. Med. 7, 6 (2024).
Kline, A. et al. Multimodal machine learning in precision health: A scoping review. npj Digit. Med. 5, 171 (2022).
Subramanian, I., Verma, S., Kumar, S., Jere, A. & Anamika, K. Multi-omics data integration, interpretation, and its application. Bioinform. Biol. Insights 14, 1177932219899051 (2020).
Amirahmadi, A., Ohlsson, M. & Etminani, K. Deep learning prediction models based on EHR trajectories: a systematic review. J. Biomed. Inform. 144, 104430 (2023).
Tong, L. et al. Integrating multi-omics data with EHR for precision medicine using advanced artificial intelligence. IEEE Rev. Biomed. Eng. 17, 80–97 (2024).
Gallagher, C. S., Ginsburg, G. S. & Musick, A. Biobanking with genetics shapes precision medicine and global health. Nat. Rev. Genet. 26, 191–202 (2025).
The All of Us Research Program Investigators The ‘All of Us’ Research Program. N. Engl. J. Med. 381, 668–676 (2019).
Bick, A. G. et al. Genomic data in the All of Us Research Program. Nature 627, 340–346 (2024).
Bycroft, C. et al. The UK Biobank resource with deep phenotyping and genomic data. Nature 562, 203–209 (2018).
Allen, N. E. et al. Prospective study design and data analysis in UK Biobank. Sci. Transl. Med. 16, eadf4428 (2024).
Kurki, M. I. et al. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature 613, 508–518 (2023).
Leitsalu, L. et al. Cohort profile: Estonian Biobank of the Estonian Genome Center, University of Tartu. Int. J. Epidemiol. 44, 1137–1147 (2015).
Nagai, A. et al. Overview of the BioBank Japan Project: study design and profile. J. Epidemiol. 27, S2–S8 (2017).
Kim, Y., Han, B.-G. & the KoGES group Cohort profile: the Korean genome and epidemiology study (KoGES) consortium. Int. J. Epidemiol. 46, e20 (2017).
Gaziano, J. M. et al. Million Veteran Program: a mega-biobank to study genetic influences on health and disease. J. Clin. Epidemiol. 70, 214–223 (2016).
McGregor, T. L. et al. Inclusion of pediatric samples in an opt-out biorepository linking DNA to de-identified medical records: pediatric BioVU. Clin. Pharmacol. Ther. 93, 204–211 (2013).
Carey, D. J. et al. The Geisinger MyCode community health initiative: an electronic health record–linked biobank for precision medicine research. Genet. Med. 18, 906–913 (2016).
Verma, A. et al. The Penn Medicine BioBank: towards a genomics-enabled learning healthcare system to accelerate precision medicine in a diverse population. J. Pers. Med. 12, 1974 (2022).
Chen, Z. et al. China Kadoorie Biobank of 0.5 million people: survey methods, baseline characteristics and long-term follow-up. Int. J. Epidemiol. 40, 1652–1666 (2011).
Cook, M. B. et al. Our Future Health: a unique global resource for discovery and translational research. Nat. Med. 31, 728–730 (2025).
Cancer Genome Atlas Research Network et al. The Cancer Genome Atlas Pan-Cancer analysis project. Nat. Genet. 45, 1113–1120 (2013).
Perez-Riverol, Y. et al. Discovering and linking public omics data sets using the Omics Discovery Index. Nat. Biotechnol. 35, 406–409 (2017).
Zhou, W. et al. Global Biobank Meta-analysis Initiative: powering genetic discovery across human disease. Cell Genom. 2, 100192 (2022).
Beesley, L. J. et al. The emerging landscape of health research based on biobanks linked to electronic health records: existing resources, statistical challenges and potential opportunities. Stat. Med. 39, 773–800 (2020).
Robinson, J. R., Wei, W.-Q., Roden, D. M. & Denny, J. C. Defining phenotypes from clinical data to drive genomic research. Annu. Rev. Biomed. Data Sci. 1, 69–92 (2018).
Ueda, D. et al. Fairness of artificial intelligence in healthcare: review and recommendations. Jpn. J. Radiol. 42, 3–15 (2024).
Al-Sahab, B., Leviton, A., Loddenkemper, T., Paneth, N. & Zhang, B. Biases in electronic health records data for generating real-world evidence: an overview. J. Healthc. Inform. Res. 8, 121–139 (2024).
Boyd, A. D. et al. Equity and bias in electronic health records data. Contemp. Clin. Trials 130, 107238 (2023).
Kachuri, L. et al. Principles and methods for transferring polygenic risk scores across global populations. Nat. Rev. Genet. 25, 8–25 (2024).
Ko, S. et al. Unsupervised discovery of ancestry-informative markers and genetic admixture proportions in biobank-scale datasets. Am. J. Hum. Genet. 110, 314–325 (2023).
Venkatesh, R. et al. Importance of genetic ancestry in pharmacogenomics for precision medicine. Pharmacogenomics 26, 747–762 (2025).
Popejoy, A. B. et al. The clinical imperative for inclusivity: race, ethnicity, and ancestry (REA) in genomics. Hum. Mutat. 39, 1713–1720 (2018).
Marees, A. T. et al. A tutorial on conducting genome-wide association studies: quality control and statistical analysis. Int. J. Methods Psychiatr. Res. 27, e1608 (2018).
Carss, K. et al. Whole-genome sequencing of 490,640 UK Biobank participants. Nature 645, 692–701 (2025).
Turner, S. et al. Quality control procedures for genome-wide association studies. Curr. Protoc. Hum. Genet. 1, 19 (2011).
McLaren, W. et al. The Ensembl Variant Effect Predictor. Genome Biol. 17, 122 (2016).
Wang, K., Li, M. & Hakonarson, H. ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res. 38, e164 (2010).
Landrum, M. J. et al. ClinVar: public archive of interpretations of clinically relevant variants. Nucleic Acids Res. 44, D862–D868 (2016).
Hamosh, A., Scott, A. F., Amberger, J. S., Bocchini, C. A. & McKusick, V. A. Online Mendelian Inheritance in Man (OMIM), a knowledgebase of human genes and genetic disorders. Nucleic Acids Res. 33, D514–D517 (2005).
Chen, S. et al. A genomic mutational constraint map using variation in 76,156 human genomes. Nature 625, 92–100 (2024).
Price, A. L. et al. Principal components analysis corrects for stratification in genome-wide association studies. Nat. Genet. 38, 904–909 (2006).
Taliun, D. et al. Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program. Nature 590, 290–299 (2021).
Auton, A. et al. A global reference for human genetic variation. Nature 526, 68–74 (2015).
Reich, D., Price, A. L. & Patterson, N. Principal component analysis of genetic data. Nat. Genet. 40, 491–492 (2008).
Zhang, D., Yin, C., Zeng, J., Yuan, X. & Zhang, P. Combining structured and unstructured data for predictive models: a deep learning approach. BMC Med. Inform. Decis. Mak. 20, 280 (2020).
O’Malley, K. J. et al. Measuring diagnoses: ICD code accuracy. Health Serv. Res. 40, 1620–1639 (2005).
Chang, E. & Mostafa, J. The use of SNOMED CT, 2013-2020: a literature review. J. Am. Med. Inform. Assoc. 28, 2017–2026 (2021).
McDonald, C. J. et al. LOINC, a universal standard for identifying laboratory observations: a 5-year update. Clin. Chem. 49, 624–633 (2003).
Liu, S., Ma, W., Moore, R., Ganesan, V. & Nelson, S. RxNorm: prescription for electronic drug information exchange. IT Prof. 7, 17–23 (2005).
Nelson, S. J., Zeng, K., Kilbourne, J., Powell, T. & Moore, R. Normalized names for clinical drugs: RxNorm at 6 years. J. Am. Med. Inform. Assoc. 18, 441–448 (2011).
CPT® code set overview. American Medical Association https://www.ama-assn.org/practice-management/cpt/cpt-code-set-overview (2026).
Reinecke, I. et al. The usage of OHDSI OMOP – a scoping review. Stud. Health Technol. Inform. 283, 95–103 (2021).
Schuemie, M. et al. Health-Analytics Data to Evidence Suite (HADES): open-source software for observational research. Stud. Health Technol. Inform. 310, 966–970 (2024).
Klann, J. G., Joss, M. A. H., Embree, K. & Murphy, S. N. Data model harmonization for the All Of Us Research Program: Transforming i2b2 data into the OMOP common data model. PLoS One 14, e0212463 (2019).
Vasilevsky, N. A. et al. Mondo: integrating disease terminology across communities. Genetics 232, iyaf215 (2026).
Robinson, P. N. et al. The Human Phenotype Ontology: a tool for annotating and analyzing human hereditary disease. Am. J. Hum. Genet. 83, 610–615 (2008).
Köhler, S. et al. The Human Phenotype Ontology in 2021. Nucleic Acids Res. 49, D1207–D1217 (2021).
Qualls, L. G. Evaluating foundational data quality in the national patient-centered clinical research network (PCORnet®). EGEMS 6, 3 (2018).
Antunes, R. S., André da Costa, C., Küderle, A., Yari, I. A. & Eskofier, B. Federated learning for healthcare: systematic review and architecture proposal. ACM Trans. Intell. Syst. Technol. 13, 54 (2022).
Hegselmann, S. et al. Large language models are powerful electronic health record encoders. npj Digit. Med. 9, 530 (2026).
Chang, C. C. et al. Second-generation PLINK: rising to the challenge of larger and richer datasets. GigaScience 4, https://doi.org/10.1186/s13742-015-0047-8 (2015).
Zhou, W. et al. Efficiently controlling for case-control imbalance and sample relatedness in large-scale genetic association studies. Nat. Genet. 50, 1335–1341 (2018).
Danecek, P. et al. Twelve years of SAMtools and BCFtools. Gigascience 10, giab008 (2021).
Ramirez, A. H. et al. The All of Us Research Program: data quality, utility, and diversity. Patterns 3, 100570 (2022).
Shih, C. C. et al. A five-safes approach to a secure and scalable genomics data repository. iScience 26, 106546 (2023).
Silva, S. et al. Fed-BioMed: A general open-source frontend framework for federated learning in healthcare. In Domain Adaptation and Representation Transfer, and Distributed and Collaborative Learning (eds Albarqouni, S. et al.) 201–210 (Springer, 2020).
Beutel, D. J. et al. Flower: a friendly federated learning research framework. Preprint at https://doi.org/10.48550/arXiv.2007.14390 (2022).
Xu, J. et al. Federated learning for healthcare informatics. J. Healthc. Inform. Res. 5, 1–19 (2021).
Silva, S. et al. Federated learning in distributed medical databases: meta-analysis of large-scale subcortical brain data. In Proc. 2019 IEEE 16th International Symposium on Biomedical Imaging (ISBI 2019) 270–274 (IEEE, 2019).
Cooray, L., Sendanayake, J., Vithanaarachchi, P. & Priyadarshana, Y. H. P. P. Deep federated learning: a systematic review of methods, applications, and challenges. Front. Comput. Sci. 7, 1617597 (2025).
Kirby, J. C. et al. PheKB: a catalog and workflow for creating electronic phenotype algorithms for transportability. J. Am. Med. Inform. Assoc. 23, 1046–1052 (2016).
Newton, K. M. et al. Validation of electronic medical record-based phenotyping algorithms: results and lessons learned from the eMERGE network. J. Am. Med. Inform. Assoc. 20, e147–e154 (2013).
Hripcsak, G. et al. Observational health data sciences and informatics (OHDSI): opportunities for observational researchers. Stud. Health Technol. Inform. 216, 574–578 (2015).
Kho, A. N. et al. Electronic medical records for genetic research: results of the eMERGE consortium. Sci. Transl. Med. 3, 79re1 (2011).
Banda, J. M., Seneviratne, M., Hernandez-Boussard, T. & Shah, N. H. Advances in electronic phenotyping: from rule-based definitions to machine learning models. Annu. Rev. Biomed. Data Sci. 1, 53–68 (2018).
Zitnik, M. et al. Machine learning for integrating data in biology and medicine: principles, practice, and opportunities. Inf. Fusion. 50, 71–91 (2019).
Ahuja, Y., Zou, Y., Verma, A., Buckeridge, D. & Li, Y. MixEHR-Guided: a guided multi-modal topic modeling approach for large-scale automatic phenotyping using the electronic health record. J. Biomed. Inform. 134, 104190 (2022).
Yang, S., Varghese, P., Stephenson, E., Tu, K. & Gronsbell, J. Machine learning approaches for electronic health records phenotyping: a methodical review. J. Am. Med. Inform. Assoc. 30, 367–381 (2023).
Luo, L. et al. PhenoTagger: a hybrid method for phenotype concept recognition using human phenotype ontology. Bioinformatics 37, 1884–1890 (2021).
Liao, K. P. et al. High-throughput multimodal automated phenotyping (MAP) with application to PheWAS. J. Am. Med. Inform. Assoc. 26, 1255–1262 (2019).
Adamson, B. et al. Approach to machine learning for extraction of real-world data variables from electronic health records. Front. Pharmacol. 14, 1180962 (2023).
Alzoubi, H. et al. A review of automatic phenotyping approaches using electronic health records. Electronics 8, 1235 (2019).
Gao, Y. & Cui, Y. Clinical time-to-event prediction enhanced by incorporating compatible related outcomes. PLoS Digit. Health 1, e0000038 (2022).
Li, Y. et al. Validation of risk prediction models applied to longitudinal electronic health record data for the prediction of major cardiovascular events in the presence of data shifts. Eur. Heart J. Digit. Health 3, 535–547 (2022).
Ramachandram, D. & Taylor, G. W. Deep multimodal learning: a survey on recent advances and trends. IEEE Signal Process. Mag. 34, 96–108 (2017).
Guo, A., Beheshti, R., Khan, Y. M., Langabeer, J. R. & Foraker, R. E. Predicting cardiovascular health trajectories in time-series electronic health records with LSTM models. BMC Med. Inform. Decis. Mak. 21, 5 (2021).
Pham, T., Tran, T., Phung, D. & Venkatesh, S. DeepCare: A deep dynamic memory model for predictive medicine. In Advances in Knowledge Discovery and Data Mining (eds Bailey, J. et al.) 30–41 (Springer, 2016).
Choi, E. et al. RETAIN: an interpretable predictive model for healthcare using reverse time attention mechanism. In Proc. 30th International Conference on Neural Information Processing Systems (eds Lee, D. D. et al.) 3512–3520 (Curran Associates Inc., 2016).
Li, Y. et al. BEHRT: transformer for electronic health records. Sci. Rep. 10, 7155 (2020).
Niu, H. et al. EHR-BERT: a BERT-based model for effective anomaly detection in electronic health records. J. Biomed. Inform. 150, 104605 (2024).
Lin, K.-W., Kuo, Y.-C., Wang, H.-Y. & Tseng, Y.-J. KAT-GNN: a knowledge-augmented temporal graph neural network for risk prediction in electronic health records. Preprint at https://doi.org/10.48550/arXiv.2511.01249 (2025).
Gao, Y. et al. Precision adverse drug reactions prediction with heterogeneous graph neural network. Adv. Sci. 12, 2404671 (2025).
Lahoti, A. et al. Mamba-3: improved sequence modeling using state space principles. Preprint at https://doi.org/10.48550/arXiv.2603.15569 (2026).
Reátegui, R. & Ratté, S. Comparison of MetaMap and cTAKES for entity extraction in clinical notes. BMC Med. Inform. Decis. Mak. 18, 74 (2018).
Savova, G. K. et al. Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES): architecture, component evaluation and applications. J. Am. Med. Inform. Assoc. 17, 507–513 (2010).
Aronson, A. R. Effective mapping of biomedical text to the UMLS Metathesaurus: the MetaMap program. In Proc. AMIA Symposium 2001 17–21 (American Medical Informatics Association, 2001).
Kang, T. et al. EliIE: an open-source information extraction system for clinical trial eligibility criteria. J. Am. Med. Inform. Assoc. 24, 1062–1071 (2017).
Lee, J. et al. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 36, 1234–1240 (2020).
Yu, X., Hu, W., Lu, S., Sun, X. & Yuan, Z. BioBERT based named entity recognition in electronic medical record. In Proc. 2019 10th International Conference on Information Technology in Medicine and Education (ITME) 49–52 (IEEE, 2019).
Wu, J. et al. A hybrid framework with large language models for rare disease phenotyping. BMC Med. Inform. Decis. Mak. 24, 289 (2024).
Alsentzer, E. et al. Publicly available clinical BERT embeddings. In Proc. 2nd Clinical Natural Language Processing Workshop (eds Rumshisky, A. et al.) 72–78 (Association for Computational Linguistics, 2019).
Yang, X. et al. A large language model for electronic health records. npj Digit. Med. 5, 194 (2022).
Khan, S. N., Danishuddin, Khan, M. W. A., Guarnera, L. & Akhtar, S. M. F. Multi-modal AI in precision medicine: integrating genomics, imaging, and EHR data for clinical insights. Front. Artif. Intell. 8, 1743921 (2026).
Tang, J., Yin, X., Lai, J., Luo, K. & Wu, D. Fusion of X-ray images and clinical data for a multimodal deep learning prediction model of osteoporosis: algorithm development and validation study. JMIR Med. Inform. 13, e70738 (2025).
Wei, H., Liu, B., Zhang, M., Shi, P. & Yuan, W. VisionCLIP: an Med-AIGC based ethical language-image foundation model for generalizable retina image analysis. Preprint at https://doi.org/10.48550/arXiv.2403.10823 (2024).
Sellergren, A. et al. MedGemma technical report. Preprint at https://doi.org/10.48550/arXiv.2507.05201 (2025).
Prottasha, M. S. I. & Rafi, N. W. MedGemma vs GPT-4: open-source and proprietary zero-shot medical disease classification from images. Preprint at https://doi.org/10.48550/arXiv.2512.23304 (2025).
Wang, L. Cross-lingual NLP: bridging language barriers with multilingual model. In Proc. 2024 International Conference on Electronics and Devices, Computational Science (ICEDCS) 1005–1012 (IEEE, 2024).
Sindane, T., Marivate, V. & Modupe, A. Cross-lingual embedding methods and applications: a systematic review for low-resourced scenarios. Nat. Lang. Proc. J. 12, 100157 (2025).
Qin, L. et al. A survey of multilingual large language models. Patterns 6, 101118 (2025).
Guo, P. et al. Steering large language models for cross-lingual information retrieval. In Proc. 47th International ACM SIGIR Conference on Research and Development in Information Retrieval 585–596 (Association for Computing Machinery, 2024).
Shah, A. Efficient cross-lingual transfer for language models. In Proc. 2025 9th International Symposium on Innovative Approaches in Smart Technologies (ISAS) 1–9 (IEEE, 2025).
Brandes, N., Linial, N. & Linial, M. PWAS: proteome-wide association study—linking genes and phenotypes by functional variation in proteins. Genome Biol. 21, 173 (2020).
Cheng, B. et al. Integrated analysis of proteome-wide and transcriptome-wide association studies identified novel genes and chemicals for vertigo. Brain Commun. 4, fcac313 (2022).
Evans, P. et al. Transcriptome-wide association studies (TWAS): methodologies, applications, and challenges. Curr. Protoc. 4, e981 (2024).
Gamazon, E. R. et al. A gene-based association method for mapping traits using reference transcriptome data. Nat. Genet. 47, 1091–1098 (2015).
Barbeira, A. N. et al. Exploiting the GTEx resources to decipher the mechanisms at GWAS loci. Genome Biol. 22, 49 (2021).
Gusev, A. et al. Integrative approaches for large-scale transcriptome-wide association studies. Nat. Genet. 48, 245–252 (2016).
Georgakis, M. K. et al. Genetically downregulated interleukin-6 signaling is associated with a favorable cardiometabolic profile. Circulation 143, 1177–1180 (2021).
Topaloudi, A. et al. PheWAS and cross-disorder analysis reveal genetic architecture, pleiotropic loci and phenotypic correlations across 11 autoimmune disorders. Front. Immunol. 14, 1147573 (2023).
Verma, A. et al. Human-disease phenotype map derived from PheWAS across 38,682 individuals. Am. J. Hum. Genet. 104, 55–64 (2019).
Ritchie, M. D., Holzinger, E. R., Li, R., Pendergrass, S. A. & Kim, D. Methods of integrating data to uncover genotype–phenotype interactions. Nat. Rev. Genet. 16, 85–97 (2015).
Žitnik, M. & Zupan, B. Data fusion by matrix factorization. IEEE Trans. Pattern Anal. Mach. Intell. 37, 41–53 (2015).
Iribarren, C. et al. Polygenic risk and incident coronary heart disease in a large multiethnic cohort. Am. J. Prev. Cardiol. 18, 100661 (2024).
Ritchie, S. C. et al. Combined clinical, metabolomic, and polygenic scores for cardiovascular risk prediction. Eur. Heart J. 47, 1861–1873 (2026).
Kim, H., Lee, G. & Chung, W. Multimodal deep learning approaches for improving polygenic risk scores with imaging data. Sci. Rep. 16, 4012 (2026).
Tong, L., Mitchel, J., Chatlin, K. & Wang, M. D. Deep learning based feature-level integration of multi-omics data for breast cancer patients survival analysis. BMC Med. Inform. Decis. Mak. 20, 225 (2020).
Farhadizadeh, M. et al. Challenges and proposed solutions in modeling multimodal data: a systematic review. Preprint at https://doi.org/10.48550/arXiv.2505.06945 (2025).
Wu, K.-H. H. et al. Integrating large scale genetic and clinical information to predict cases of heart failure. Commun. Med. 5, 493 (2025).
Mataraso, S. J. et al. A machine learning approach to leveraging electronic health records for enhanced omics analysis. Nat. Mach. Intell. 7, 293–306 (2025).
Eijpe, A. et al. Disentangled and interpretable multimodal attention fusion for cancer survival prediction. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2025 (eds Gee, J. C. et al.) 117–127 (Springer, 2025).
Lin, J. et al. Integration of biomarker polygenic risk score improves prediction of coronary heart disease. JACC Basic Transl. Sci. 8, 1489–1499 (2023).
Türkmen, D. et al. Polygenic scores for cardiovascular risk factors improve estimation of clinical outcomes in CCB treatment compared to pharmacogenetic variants alone. Pharmacogenomics J. 24, 12 (2024).
Liu, Y., Huse, J. & Kannan, K. Expression graph network framework for biomarker discovery. Brief. Bioinform. 26, bbaf559 (2025).
Wang, T. et al. MOGONET integrates multi-omics data using graph convolutional networks allowing patient classification and biomarker identification. Nat. Commun. 12, 3445 (2021).
Amar, J. et al. Integrating genomics into multimodal EHR foundation models. Preprint at https://doi.org/10.48550/arXiv.2510.23639 (2025).
Nguyen, E. et al. Sequence modeling and design from molecular to genome scale with Evo. Science 386, eado9336 (2024).
Brixi, G. et al. Genome modelling and design across all domains of life with Evo 2. Nature 652, 1349–1361 (2026).
Dalla-Torre, H. et al. Nucleotide Transformer: building and evaluating robust foundation models for human genomics. Nat. Methods 22, 287–297 (2025).
Nguyen, E. et al. HyenaDNA: long-range genomic sequence modeling at single nucleotide resolution. In Proc. 37th International Conference on Neural Information Processing Systems (eds Oh, A. et al.) 43177–43201 (Curran Associates Inc., 2023).
Cui, H. et al. scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nat. Methods 21, 1470–1480 (2024).
Hao, M. et al. Large-scale foundation model on single-cell transcriptomics. Nat. Methods 21, 1481–1491 (2024).
Steinberg, E. et al. Language models are an effective representation learning technique for electronic health record data. J. Biomed. Inform. 113, 103637 (2021).
Wornow, M., Thapa, R., Steinberg, E., Fries, J. A. & Shah, N. H. EHRSHOT: an EHR benchmark for few-shot evaluation of foundation models. In Advances in Neural Information Processing Systems (eds Oh, A. et al.) 67125–67137 (Curran Associates, Inc., 2023).
Renc, P. et al. Zero shot health trajectory prediction using transformer. npj Digit. Med. 7, 256 (2024).
Steinberg, E., Fries, J., Xu, Y. & Shah, N. MOTOR: a time-to-event foundation model for structured medical records. In International Conference on Learning Representations (eds Kim, B. et al.) 41038–41077 (ICLR, 2024).
Zhao, T. et al. BiomedParse: a biomedical foundation model for image parsing of everything everywhere all at once. Nat. Methods 22, 166–176 (2025).
Wu, C., Zhang, X., Zhang, Y., Wang, Y. & Xie, W. Towards generalist foundation model for radiology by leveraging web-scale 2D&3D medical data. Nat Commun. 16, 7866 (2025).
Cardoso, M. J. et al. MONAI: an open-source framework for deep learning in healthcare. Preprint at https://doi.org/10.48550/arXiv.2211.02701 (2022).
Yang, L. et al. Advancing multimodal medical capabilities of Gemini. Preprint at https://doi.org/10.48550/arXiv.2405.03162 (2024).
Ahlqvist, E. et al. Novel subgroups of adult-onset diabetes and their association with outcomes: a data-driven cluster analysis of six variables. Lancet Diabetes Endocrinol. 6, 361–369 (2018).
Qiu, J. et al. Deep representation learning for clustering longitudinal survival data from electronic health records. Nat. Commun. 16, 2534 (2025).
Soman, K. et al. Early detection of Parkinson’s disease through enriching the electronic health record using a biomedical knowledge graph. Front. Med. 10, 1081087 (2023).
Li, C. et al. Improving cardiovascular risk prediction through machine learning modelling of irregularly repeated electronic health records. Eur. Heart J. Digit. Health 5, 30–40 (2024).
Shen, L. et al. Genetic analysis of quantitative phenotypes in AD and MCI: imaging, cognition and biomarkers. Brain Imaging Behav. 8, 183–207 (2014).
Crane, P. K. et al. Cognitively defined Alzheimer’s dementia subgroups have distinct atrophy patterns. Alzheimers Dement. 20, 1739–1752 (2024).
Ferraro, P. M. et al. Clinical and biological underpinnings of longitudinal atrophy pattern progression in Alzheimer’s disease. J. Alzheimer’s Dis. 103, 243–255 (2025).
Higginbotham, L. et al. Unbiased classification of the elderly human brain proteome resolves distinct clinical and pathophysiological subtypes of cognitive impairment. Neurobiology of Disease 186, 106286 https://doi.org/10.1016/j.nbd.2023.106286 (2023).
Hähnel, T. et al. Progression subtypes in Parkinson’s disease identified by a data-driven multi cohort analysis. npj Parkinsons Dis. 10, 95 (2024).
Krix, S. et al. MultiGML: multimodal graph machine learning for prediction of adverse drug events. Heliyon 9, e19441 (2023).
Adam, G. et al. Machine learning approaches to drug response prediction: challenges and recent progress. npj Precis. Oncol. 4, 19 (2020).
Ozery-Flato, M., Goldschmidt, Y., Shaham, O., Ravid, S. & Yanover, C. Framework for identifying drug repurposing candidates from observational healthcare data. JAMIA Open 3, 536–544 (2020).
Hernán, M. A., Wang, W. & Leaf, D. E. Target trial emulation: a framework for causal inference from observational data. JAMA 328, 2446–2447 (2022).
Kaufman, H., Rappoport, N., Gilad, A. & Linial, M. Advancing causal inference in medicine using biobank data. J. Biomed. Inform. 171, 104903 (2025).
Yao, M., Wang, A., Li, X. & Liu, Z. Mendelian randomization methods for causal inference: estimands, identification and inference. Stat. Med. 45, e70394 (2026).
Al Khzem, A. H. & Wali, S. M. Drug repurposing as an effective drug discovery strategy: a critical review. Drug Des. Dev. Ther. 19, 12019–12034 (2025).
Lau-Min, K. S. et al. Impact of integrating genomic data into the electronic health record on genetics care delivery. Genet. Med. 24, 2338–2350 (2022).
Rasmussen-Torvik, L. J. et al. Design and anticipated outcomes of the eMERGE-PGx project: a multi-center pilot for pre-emptive pharmacogenomics in electronic health record systems. Clin. Pharmacol. Ther. 96, 482–489 (2014).
Murugan, M. et al. Genomic considerations for FHIR®; eMERGE implementation lessons. J. Biomed. Inform. 118, 103795 (2021).
Cheng, D. T. et al. Memorial Sloan Kettering-integrated mutation profiling of actionable cancer targets (MSK-IMPACT): a hybridization capture-based next-generation sequencing clinical assay for solid tumor molecular oncology. J. Mol. Diagn. 17, 251–264 (2015).
Klein, H. et al. MatchMiner: an open-source platform for cancer precision medicine. npj Precis. Oncol. 6, 69 (2022).
Vassy, J. L. et al. Genomic risk model to implement precision prostate cancer screening in clinical care: the ProGRESS study. Nat. Cancer 7, 352–367 (2026).
Geisinger. Geisinger launches ‘MyCode-Connect’: a new era of precision health integrating real-world data and multi-omics. Geisinger News Releases https://www.geisinger.org/about-geisinger/news-and-media/news-releases/2026/03/10/13/54/geisinger-launches-mycode-connect-a-new-era-of-precision-health (2026).
Iribarren, C. et al. Abstract 4355586: Enhancing the prevent equation with a polygenic risk score: clinical utility evaluation. Circulation 152, A4355586 (2025).
Rockowitz, S. et al. Children’s rare disease cohorts: an integrative research and clinical genomics initiative. npj Genom. Med. 5, 29 (2020).
Martin, A. R. et al. Human demographic history impacts genetic risk prediction across diverse populations. Am. J. Hum. Genet. 100, 635–649 (2017).
Linder, J. E. et al. Returning integrated genomic risk and clinical recommendations: the eMERGE study. Genet. Med. 25, 100006 (2023).
Dolin, R. H., Boxwala, A. & Shalaby, J. A pharmacogenomics clinical decision support service based on FHIR and CDS Hooks. Methods Inf. Med. 57, e115–e123 (2018).
Mandel, J. C., Kreda, D. A., Mandl, K. D., Kohane, I. S. & Ramoni, R. B. SMART on FHIR: a standards-based, interoperable apps platform for electronic health records. J. Am. Med. Inform. Assoc. 23, 899–908 (2016).
Adnan, M., Kalra, S., Cresswell, J. C., Taylor, G. W. & Tizhoosh, H. R. Federated learning and differential privacy for medical image analysis. Sci. Rep. 12, 1953 (2022).
Ramos, E. et al. Pharmacogenomics, ancestry and clinical decision making for global populations. Pharmacogenomics J. 14, 217–222 (2014).
Huang, R. et al. Evaluation and bias analysis of large language models in generating synthetic electronic health records: comparative study. J. Med. Internet Res. 27, e65317 (2025).
Kullo, I. J. et al. Polygenic scores in biomedical research. Nat. Rev. Genet. 23, 524–532 (2022).
Abràmoff, M. D. et al. Considerations for addressing bias in artificial intelligence for health equity. npj Digit. Med. 6, 170 (2023).
Azad, T. D., Krumholz, H. M. & Saria, S. Principles to guide clinical AI readiness and move from benchmarks to real-world evaluation. Nat. Med. 32, 802–804 (2026).
Hassija, V. et al. Interpreting black-box models: a review on explainable artificial intelligence. Cogn. Comput. 16, 45–74 (2024).
Loh, H. W. et al. Application of explainable artificial intelligence for healthcare: a systematic review of the last decade (2011–2022). Comput. Methods Programs Biomed. 226, 107161 (2022).
Linardatos, P., Papastefanopoulos, V. & Kotsiantis, S. Explainable AI: a review of machine learning interpretability methods. Entropy 23, 18 (2021).
Watson, J. et al. Overcoming barriers to the adoption and implementation of predictive modeling and machine learning in clinical care: what can we learn from US academic medical centers? JAMIA Open 3, 167–172 (2020).
Singh, V., Cheng, S., Kwan, A. C. & Ebinger, J. United States food and drug administration regulation of clinical software in the era of artificial intelligence and machine learning. Mayo Clin. Proc. Digit. Health 3, 100231 (2025).
Atmaca, U. I. et al. Data-driven medical devices and the EU MDR: mapping gaps in standards for regulatory compliance. npj Health Syst. 3, 21 (2026).
Meszaros, J., Minari, J. & Huys, I. The future regulation of artificial intelligence systems in healthcare services and medical research in the European Union. Front. Genet. 13, 927721 (2022).
Karunanayake, N. Next-generation agentic AI for transforming healthcare. Inform. Health 2, 73–83 (2025).
Liu, F. et al. A foundational architecture for AI agents in healthcare. Cell Rep. Med. 6, 102374 (2025).
Waight, M. C. et al. Personalized heart digital twins detect substrate abnormalities in scar-dependent ventricular tachycardia. Circulation 151, 521–533 (2025).
Pan, R., Sun, H., Chen, X., Pedrielli, G. & Huang, J. Human digital twin: data, models, applications, and challenges. Preprint at https://doi.org/10.48550/arXiv.2508.13138 (2025).
DeMeo, B. et al. Active learning framework leveraging transcriptomics identifies modulators of disease phenotypes. Science 390, eadi8577 (2025).
Boshar, S. et al. A foundational model for joint sequence-function multi-species modeling at scale for long-range genomic prediction. Preprint at bioRxiv https://doi.org/10.64898/2025.12.22.695963 (2025).
Shen, T. et al. Accurate RNA 3D structure prediction using a language model-based deep learning approach. Nat. Methods 21, 2287–2298 (2024).
Chen, J. et al. Interpretable RNA foundation model from unannotated data for highly accurate RNA structure and function predictions. Preprint at https://doi.org/10.48550/arXiv.2204.00300 (2022).
Peng, C. et al. A study of generative large language model for medical research and healthcare. npj Digit. Med. 6, 210 (2023).
Inouye, M. et al. Genomic risk prediction of coronary artery disease in 480,000 adults. J. Am. Coll. Cardiol. 72, 1883–1893 (2018).
Verma, S. S. et al. Evaluating the frequency and the impact of pharmacogenetic alleles in an ancestrally diverse Biobank population. J. Transl. Med. 20, 550 (2022).
Zhu, J., Liu, H., Liu, X., Chen, C. & Shu, M. Cardiovascular disease detection based on deep learning and multi-modal data fusion. Biomed. Signal Process. Control 99, 106882 (2025).
Zhang, X. et al. Data-driven subtyping of Parkinson’s disease using longitudinal clinical records: a cohort study. Sci. Rep. 9, 797 (2019).
Nelson, C. A., Bove, R., Butte, A. J. & Baranzini, S. E. Embedding electronic health records onto a knowledge network recognizes prodromal features of multiple sclerosis and predicts diagnosis. J. Am. Med. Inform. Assoc. 29, 424–434 (2022).
Jin, W., Xu, Y. & Wang, Z. Modeling Alzheimer’s disease biomarkers’ trajectory in the absence of a gold standard using a Bayesian approach. Stat. Med. 44, e70283 (2025).
Lachmann, M. et al. Subphenotyping of patients with aortic stenosis by unsupervised agglomerative clustering of echocardiographic and hemodynamic data. JACC Cardiovasc. Interv. 14, 2127–2140 (2021).
Chaudhary, K., Poirion, O. B., Lu, L. & Garmire, L. X. Deep learning-based multi-omics integration robustly predicts survival in liver cancer. Clin. Cancer Res. 24, 1248–1259 (2018).
Schlosser, P. et al. Transcriptome- and proteome-wide association studies nominate determinants of kidney function and damage. Genome Biol. 24, 150 (2023).
Sun, B. B. et al. Plasma proteomic associations with genetics and health in the UK Biobank. Nature 622, 329–338 (2023).
Himmelstein, D. S. et al. Systematic integration of biomedical knowledge prioritizes drugs for repurposing. eLife 6, e26726 (2017).
Minikel, E. V., Painter, J. L., Dong, C. C. & Nelson, M. R. Refining the impact of genetic evidence on clinical success. Nature 629, 624–629 (2024).
Gao, C. et al. Proteome-wide association study for finding druggable targets in progression and onset of Parkinson’s disease. CNS Neurosci. Ther. 31, e70294 (2025).
Zitnik, M., Agrawal, M. & Leskovec, J. Modeling polypharmacy side effects with graph convolutional networks. Bioinformatics 34, i457–i466 (2018).
Miotto, R., Li, L., Kidd, B. A. & Dudley, J. T. Deep Patient: an unsupervised representation to predict the future of patients from the electronic health records. Sci. Rep. 6, 26094 (2016).
Estiri, H. et al. Transitive sequencing medical records for mining predictive and interpretable temporal representations. Patterns 1, 100051 (2020).
Li, D., Xing, W., Zhao, J., Shi, C. & Wang, F. Multimodal deep learning for predicting in-hospital mortality in heart failure patients using longitudinal chest X-rays and electronic health records. Int. J. Cardiovasc. Imaging 41, 427–440 (2025).
Barr, P. B. et al. Polygenic risk factors for comorbid diagnoses in individuals with substance use disorders: a phenome-wide survival analysis. Psychol. Med. 56, e174 (2026).
Forero, D. A. et al. Current needs for human and medical genomics research infrastructure in low and middle income countries. J. Med. Genet. 53, 438–440 (2016).
Woldemariam, M. T. & Jimma, W. Adoption of electronic health record systems to enhance the quality of healthcare in low-income countries: a systematic review. BMJ Health Care Inform. 30, e100704 (2023).
The H3Africa Consortium et al. Enabling the genomic revolution in Africa. Science 344, 1346–1348 (2014).
Were, M. C. et al. mUzima mobile electronic health record (EHR) system: development and implementation at scale. J. Med. Internet Res. 23, e26381 (2021).
Robbiati, C. et al. Improving TB surveillance and patients’ quality of care through improved data collection in Angola: development of an electronic medical record system in two health facilities of Luanda. Front. Public Health 10, 745928 (2022).