• Hoot, N. R. & Aronsky, D. Systematic review of emergency department crowding: causes, effects, and solutions. Ann. Emerg. Med. 52, 126–136 (2008).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Morley, C., Unwin, M., Peterson, G. M., Stankovich, J. & Kinsman, L. Emergency department crowding: a systematic review of causes, consequences and solutions. PLoS ONE 13, e0203316 (2018).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Voaklander, B. et al. Interventions to improve consultations in the emergency department: a systematic review. Acad. Emerg. Med. 29, 1475–1495 (2022).

    Article 
    PubMed 

    Google Scholar
     

  • Brick, C. et al. The impact of consultation on length of stay in tertiary care emergency departments. Emerg. Med. J. 31, 134–138 (2014).

    Article 
    PubMed 

    Google Scholar
     

  • Piliuk, K. & Tomforde, S. Artificial intelligence in emergency medicine: a systematic literature review. Int. J. Med. Inform. 180, 105274 (2023).

    Article 
    PubMed 

    Google Scholar
     

  • Grant, K., McParland, A., Mehta, S. & Ackery, A. D. Artificial intelligence in emergency medicine: surmountable barriers with revolutionary potential. Ann. Emerg. Med. 75, 721–726 (2020).

    Article 
    PubMed 

    Google Scholar
     

  • Nagendran, M. et al. Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies. BMJ 368, m689 (2020).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Plana, D. et al. Randomized clinical trials of machine learning interventions in health care: a systematic review. JAMA Netw. Open 5, e2233946 (2022).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Han, T. et al. Randomised controlled trials evaluating artificial intelligence in clinical practice: a scoping review. Lancet Digit. Health 6, e367–e373 (2024).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Oikonomidi, T. et al. Applications of artificial intelligence for real-world evidence generation: a protocol for a living scoping review. BMJ Open 16, e109725 (2026).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Kelly, C. J., Karthikesalingam, A., Suleyman, M., Corrado, G. & King, D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 17, 195 (2019).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Wiens, J. et al. Do no harm: a roadmap for responsible machine learning for health care. Nat. Med. 25, 1337–1340 (2019).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  • Vasey, B. et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat. Med. 28, 924–933 (2022).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  • Thirunavukarasu, A. J. et al. Large language models in medicine. Nat. Med. 29, 1930–1940 (2023).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  • Singhal, K. et al. Large language models encode clinical knowledge. Nature 620, 172–180 (2023).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Jiang, L. Y. et al. Health system-scale language models are all-purpose prediction engines. Nature 619, 357–362 (2023).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Williams, C. Y. K., Miao, B. Y., Kornblith, A. E. & Butte, A. J. Evaluating the use of large language models to provide clinical recommendations in the emergency department. Nat. Commun. 15, 8236 (2024).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Hager, P. et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat. Med. 30, 2613–2622 (2024).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Preiksaitis, C. & Rose, C. The role of large language models in transforming emergency medicine: scoping review. JMIR Med. Inform. 12, e53787 (2024).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Gilboy, N., Tanabe, T., Travers, D. & Rosenau, A. M.Emergency Severity Index (ESI): A Triage Tool for Emergency Department Care, Version 4. AHRQ Publication No. 12-0014 (AHRQ, 2011).

  • Greenhalgh, T. et al. Beyond adoption: a new framework for theorizing and evaluating nonadoption, abandonment, and challenges to the scale-up, spread, and sustainability of health and care technologies. J. Med. Internet Res. 19, e367 (2017).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • R Core Team R: A Language and Environment for Statistical Computing (R Core Team, 2024).

  • Ancker, J. S. et al. Effects of workload, work complexity, and repeated alerts on alert fatigue in a clinical decision support system. BMC Med. Inform. Decis. Mak. 17, 36 (2017).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Kawamoto, K., Houlihan, C. A., Balas, E. A. & Lobach, D. F. Improving clinical practice using clinical decision support systems: a systematic review of trials to identify features critical to success. BMJ 330, 765 (2005).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Roshanov, P. S. et al. Features of effective computerised clinical decision support systems: meta-regression of 162 randomised trials. BMJ 346, f657 (2013).

    Article 
    PubMed 

    Google Scholar
     

  • Proctor, E. K. et al. Outcomes for implementation research: conceptual distinctions, measurement challenges, and research agenda. Adm. Policy Ment. Health 38, 65–76 (2011).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Angrist, J. D., Imbens, G. W. & Rubin, D. B. Identification of causal effects using instrumental variables. J. Am. Stat. Assoc. 91, 444–455 (1996).

    Article 

    Google Scholar
     

  • Kwong, J. C. C., Nickel, G. C., Wang, S. C. Y. & Kvedar, J. C. Integrating artificial intelligence into healthcare systems: more than just the algorithm. npj Digit. Med. 7, 52 (2024).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • McCambridge, J., Witton, J. & Elbourne, D. R. Systematic review of the Hawthorne effect: new concepts are needed to study research participation effects. J. Clin. Epidemiol. 67, 267–277 (2014).

    Article 
    PubMed 

    Google Scholar
     

  • Geskey, J. M., Geeting, G., West, C. & Hollenbeak, C. S. Improved physician consult response times in an academic emergency department after implementation of an institutional guideline. J. Emerg. Med. 44, 999–1006 (2013).

    Article 
    PubMed 

    Google Scholar
     

  • Soong, C., High, S., Morgan, M. W. & Ovens, H. A novel approach to improving emergency department consultant response times. BMJ Qual. Saf. 22, 299–305 (2013).

    Article 
    PubMed 

    Google Scholar
     

  • VanderWeele, T. J. & Ding, P. Sensitivity analysis in observational research: introducing the E-value. Ann. Intern. Med. 167, 268–274 (2017).

    Article 
    PubMed 

    Google Scholar
     

  • Donabedian, A. The quality of care: how can it be assessed?. JAMA 260, 1743–1748 (1988).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  • Mant, J. Process versus outcome indicators in the assessment of quality of health care. Int. J. Qual. Health Care 13, 475–480 (2001).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  • Torgerson, D. J. Contamination in trials: is cluster randomisation the answer? BMJ 322, 355–357 (2001).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Hemming, K., Haines, T. P., Chilton, P. J., Girling, A. J. & Lilford, R. J. The stepped wedge cluster randomised trial: rationale, design, analysis, and reporting. BMJ 350, h391 (2015).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  • Lewis P. et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Proc. 34th International Conference on Neural Information Processing Systems (NIPS ’20) (eds Larochelle, H. et al.) 9459–9474 (Curran, 2020).

  • Topol, E. J. High-performance medicine: the convergence of human and artificial intelligence. Nat. Med. 25, 44–56 (2019).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  • Ethics and Governance of Artificial Intelligence for Health: WHO Guidance (WHO, 2021).

  • Shin, S., Lee, S. H., Kim, D. H. & Kim, K. H. The impact of the improvement in internal medicine consultation process on emergency department length of stay. Am. J. Emerg. Med. 36, 620–624 (2018).

    Article 
    PubMed 

    Google Scholar
     

  • Ravi, A., Shochat, G., Wang, R. C. & Khanna, R. Improvements to emergency department length of stay and user satisfaction after implementation of an integrated consult order. J. Am. Coll. Emerg. Physicians Open 4, e12922 (2023).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Austin, P. C. Balance diagnostics for comparing the distribution of baseline covariates between treatment groups in propensity-score matched samples. Stat. Med. 28, 3083–3107 (2009).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Greenland, S. An introduction to instrumental variables for epidemiologists. Int. J. Epidemiol. 29, 722–729 (2000).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  • Staiger, D. & Stock, J. H. Instrumental variables regression with weak instruments. Econometrica 65, 557–586 (1997).

    Article 

    Google Scholar
     

  • Hodges, J. L. Jr & Lehmann, E. L. Estimates of location based on rank tests. Ann. Math. Stat. 34, 598–611 (1963).

    Article 

    Google Scholar
     

  • Meunier, P. Y., Raynaud, C., Guimaraes, E., Gueyffier, F. & Letrilliart, L. Barriers and facilitators to the use of clinical decision support systems in primary care: a mixed-methods systematic review. Ann. Fam. Med. 21, 57–69 (2023).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  • Python v.3.13 (Python Software Foundation, 2024).

  • llironlibo. llironlibo/SHAKED-analysis: v1.1.0. Zenodo https://doi.org/10.5281/zenodo.20736931 (2026).