Hoot, N. R. & Aronsky, D. Systematic review of emergency department crowding: causes, effects, and solutions. Ann. Emerg. Med. 52, 126–136 (2008).
Morley, C., Unwin, M., Peterson, G. M., Stankovich, J. & Kinsman, L. Emergency department crowding: a systematic review of causes, consequences and solutions. PLoS ONE 13, e0203316 (2018).
Voaklander, B. et al. Interventions to improve consultations in the emergency department: a systematic review. Acad. Emerg. Med. 29, 1475–1495 (2022).
Brick, C. et al. The impact of consultation on length of stay in tertiary care emergency departments. Emerg. Med. J. 31, 134–138 (2014).
Piliuk, K. & Tomforde, S. Artificial intelligence in emergency medicine: a systematic literature review. Int. J. Med. Inform. 180, 105274 (2023).
Grant, K., McParland, A., Mehta, S. & Ackery, A. D. Artificial intelligence in emergency medicine: surmountable barriers with revolutionary potential. Ann. Emerg. Med. 75, 721–726 (2020).
Nagendran, M. et al. Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies. BMJ 368, m689 (2020).
Plana, D. et al. Randomized clinical trials of machine learning interventions in health care: a systematic review. JAMA Netw. Open 5, e2233946 (2022).
Han, T. et al. Randomised controlled trials evaluating artificial intelligence in clinical practice: a scoping review. Lancet Digit. Health 6, e367–e373 (2024).
Oikonomidi, T. et al. Applications of artificial intelligence for real-world evidence generation: a protocol for a living scoping review. BMJ Open 16, e109725 (2026).
Kelly, C. J., Karthikesalingam, A., Suleyman, M., Corrado, G. & King, D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 17, 195 (2019).
Wiens, J. et al. Do no harm: a roadmap for responsible machine learning for health care. Nat. Med. 25, 1337–1340 (2019).
Vasey, B. et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat. Med. 28, 924–933 (2022).
Thirunavukarasu, A. J. et al. Large language models in medicine. Nat. Med. 29, 1930–1940 (2023).
Singhal, K. et al. Large language models encode clinical knowledge. Nature 620, 172–180 (2023).
Jiang, L. Y. et al. Health system-scale language models are all-purpose prediction engines. Nature 619, 357–362 (2023).
Williams, C. Y. K., Miao, B. Y., Kornblith, A. E. & Butte, A. J. Evaluating the use of large language models to provide clinical recommendations in the emergency department. Nat. Commun. 15, 8236 (2024).
Hager, P. et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat. Med. 30, 2613–2622 (2024).
Preiksaitis, C. & Rose, C. The role of large language models in transforming emergency medicine: scoping review. JMIR Med. Inform. 12, e53787 (2024).
Gilboy, N., Tanabe, T., Travers, D. & Rosenau, A. M.Emergency Severity Index (ESI): A Triage Tool for Emergency Department Care, Version 4. AHRQ Publication No. 12-0014 (AHRQ, 2011).
Greenhalgh, T. et al. Beyond adoption: a new framework for theorizing and evaluating nonadoption, abandonment, and challenges to the scale-up, spread, and sustainability of health and care technologies. J. Med. Internet Res. 19, e367 (2017).
R Core Team R: A Language and Environment for Statistical Computing (R Core Team, 2024).
Ancker, J. S. et al. Effects of workload, work complexity, and repeated alerts on alert fatigue in a clinical decision support system. BMC Med. Inform. Decis. Mak. 17, 36 (2017).
Kawamoto, K., Houlihan, C. A., Balas, E. A. & Lobach, D. F. Improving clinical practice using clinical decision support systems: a systematic review of trials to identify features critical to success. BMJ 330, 765 (2005).
Roshanov, P. S. et al. Features of effective computerised clinical decision support systems: meta-regression of 162 randomised trials. BMJ 346, f657 (2013).
Proctor, E. K. et al. Outcomes for implementation research: conceptual distinctions, measurement challenges, and research agenda. Adm. Policy Ment. Health 38, 65–76 (2011).
Angrist, J. D., Imbens, G. W. & Rubin, D. B. Identification of causal effects using instrumental variables. J. Am. Stat. Assoc. 91, 444–455 (1996).
Kwong, J. C. C., Nickel, G. C., Wang, S. C. Y. & Kvedar, J. C. Integrating artificial intelligence into healthcare systems: more than just the algorithm. npj Digit. Med. 7, 52 (2024).
McCambridge, J., Witton, J. & Elbourne, D. R. Systematic review of the Hawthorne effect: new concepts are needed to study research participation effects. J. Clin. Epidemiol. 67, 267–277 (2014).
Geskey, J. M., Geeting, G., West, C. & Hollenbeak, C. S. Improved physician consult response times in an academic emergency department after implementation of an institutional guideline. J. Emerg. Med. 44, 999–1006 (2013).
Soong, C., High, S., Morgan, M. W. & Ovens, H. A novel approach to improving emergency department consultant response times. BMJ Qual. Saf. 22, 299–305 (2013).
VanderWeele, T. J. & Ding, P. Sensitivity analysis in observational research: introducing the E-value. Ann. Intern. Med. 167, 268–274 (2017).
Donabedian, A. The quality of care: how can it be assessed?. JAMA 260, 1743–1748 (1988).
Mant, J. Process versus outcome indicators in the assessment of quality of health care. Int. J. Qual. Health Care 13, 475–480 (2001).
Torgerson, D. J. Contamination in trials: is cluster randomisation the answer? BMJ 322, 355–357 (2001).
Hemming, K., Haines, T. P., Chilton, P. J., Girling, A. J. & Lilford, R. J. The stepped wedge cluster randomised trial: rationale, design, analysis, and reporting. BMJ 350, h391 (2015).
Lewis P. et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Proc. 34th International Conference on Neural Information Processing Systems (NIPS ’20) (eds Larochelle, H. et al.) 9459–9474 (Curran, 2020).
Topol, E. J. High-performance medicine: the convergence of human and artificial intelligence. Nat. Med. 25, 44–56 (2019).
Ethics and Governance of Artificial Intelligence for Health: WHO Guidance (WHO, 2021).
Shin, S., Lee, S. H., Kim, D. H. & Kim, K. H. The impact of the improvement in internal medicine consultation process on emergency department length of stay. Am. J. Emerg. Med. 36, 620–624 (2018).
Ravi, A., Shochat, G., Wang, R. C. & Khanna, R. Improvements to emergency department length of stay and user satisfaction after implementation of an integrated consult order. J. Am. Coll. Emerg. Physicians Open 4, e12922 (2023).
Austin, P. C. Balance diagnostics for comparing the distribution of baseline covariates between treatment groups in propensity-score matched samples. Stat. Med. 28, 3083–3107 (2009).
Greenland, S. An introduction to instrumental variables for epidemiologists. Int. J. Epidemiol. 29, 722–729 (2000).
Staiger, D. & Stock, J. H. Instrumental variables regression with weak instruments. Econometrica 65, 557–586 (1997).
Hodges, J. L. Jr & Lehmann, E. L. Estimates of location based on rank tests. Ann. Math. Stat. 34, 598–611 (1963).
Meunier, P. Y., Raynaud, C., Guimaraes, E., Gueyffier, F. & Letrilliart, L. Barriers and facilitators to the use of clinical decision support systems in primary care: a mixed-methods systematic review. Ann. Fam. Med. 21, 57–69 (2023).
Python v.3.13 (Python Software Foundation, 2024).
llironlibo. llironlibo/SHAKED-analysis: v1.1.0. Zenodo https://doi.org/10.5281/zenodo.20736931 (2026).