{"id":61522,"date":"2026-06-04T01:46:06","date_gmt":"2026-06-04T01:46:06","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/61522\/"},"modified":"2026-06-04T01:46:06","modified_gmt":"2026-06-04T01:46:06","slug":"keeping-a-healthy-degree-of-ai-skepticism-understanding-ai-metrics-in-sleep-medicine","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/61522\/","title":{"rendered":"Keeping a Healthy Degree of AI Skepticism: Understanding AI Metrics in Sleep Medicine"},"content":{"rendered":"<p>By Daniel Rongo, MD; Scott Ryals, MD; Mattina A. Davenport, PhD; and Trung Le, PhD, on behalf of the AASM Artificial Intelligence in Sleep Medicine Committee <\/p>\n<p>As artificial intelligence (AI) becomes integrated into the clinical workflow of sleep medicine, the clinician\u2019s role is evolving. Clinicians should understand how AI models function and decide whether they perform as intended. Advanced expertise in computer science is not required to advocate for our patients. A working knowledge of the AI lifecycle, performance metrics, and common statistical pitfalls is sufficient to prevent being misled by impressive-looking numbers. These skills help to ensure AI improves care as the field of sleep medicine enters a new era.<\/p>\n<p>The AI model lifecycle<\/p>\n<p>Every AI model begins with defining a problem or objective. Relevant data are collected and preprocessed so that the model can learn patterns and predict outcomes (i.e., supervised learning). The model is then evaluated on separate test data not used during training (i.e., for testing\/validating data). Demonstrating generalizability across diverse patient populations, institutions, and clinical settings is essential before deployment. Deployment is not the endpoint. Continuous monitoring is required to detect performance drift, bias, and unintended clinical consequences. One loop of this iterative life cycle is demonstrated in Figure 1. Ongoing oversight by scientists, administrators, and clinicians is necessary to maintain AI model safety and usefulness.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" width=\"718\" height=\"384\" title=\"figure1_AImodel\" data-lazy-type=\"image\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/figure1_AImodel.png\" alt=\"\" class=\"lazy lazy-hidden img-responsive wp-image-88300\"  \/><\/p>\n<p>Figure 1: The AI model lifecycle<\/p>\n<p>Managing metrics of performance<\/p>\n<p>AI models optimize predictions based on a chosen objective. The challenge is deciding which performance metrics matter most clinically. Consider predicting a rare disease (e.g., narcolepsy) with a prevalence of 1%. Accuracy alone is meaningless in rare conditions. In this setting, minimizing false negatives is critical. The preferred metric becomes recall (sensitivity): the ability to detect true cases. Model training involves optimizing toward specific performance targets. However, optimizing one metric introduces tradeoffs. A model tuned for recall may increase false positives. Selecting the correct metric requires aligning performance goals with clinical risk.<\/p>\n<p>A major concern is overfitting, when a model memorizes noise in the training patient data and fails on new patients. True validation requires testing on unseen, real-world data. Strong performance during development does not guarantee safe clinical performance.<\/p>\n<p>Different metrics for different models<\/p>\n<p>No single metric applies to every task. The appropriate metric depends on the clinical question. Table 1 summarizes common sleep medicine applications and the metrics most relevant to each. These are not exhaustive but provide a framework for evaluating claims about AI model performance.<\/p>\n<p>Table 1:\u00a0Examples of tasks\u00a0for sleep medicine and the metrics that matter\u00a0<\/p>\n<p>Type of Model<br \/>\nExample (Sleep Medicine)<br \/>\nMetrics<\/p>\n<p>Classification<br \/>\nPredicting epoch sleep stage or CPAP noncompliance<br \/>\nAccuracy, precision, recall, F1 score, area under the curve of a receiver operating characteristic (AUC ROC), Cohen\u2019s kappa, Matthews correlation coefficient (MCC), calibration<\/p>\n<p>Regression<br \/>\nPredicting apnea\u2013hypopnea index (AHI) or total sleep time from physiologic or wearable data<br \/>\nMean absolute error (MAE), root mean square error (RMSE), R2, Pearson r, Bland\u2013Altman bias<\/p>\n<p>Time-series \/ signal processing<br \/>\nDetecting respiratory events or sleep stages (K-complex, sleep spindle, frequencies) from raw EEG\/airflow signals using Fourier transform, wavelet transform, recurrent neural networks (RNNs), or transformers<br \/>\nEvent-level sensitivity\/specificity, AUC ROC, Cohen\u2019s kappa vs. expert scoring, correlation coefficients, per-epoch accuracy<\/p>\n<p>Causal inference models<br \/>\nEstimating the effect of PAP therapy on daytime sleepiness after controlling confounders<br \/>\nAverage treatment effect (ATE), conditional ATE (CATE), overlap diagnostics, balance metrics<\/p>\n<p>Generative AI \/ large language models<br \/>\nAI scribe systems for note writing<br \/>\nHallucination rate, omission rate, attribution error rate, critical error rate, note acceptance rate, median time saved per note<\/p>\n<p>Classification tasks, such as sleep staging or predicting CPAP adherence, are especially common in clinical decision-making. Table 2 describes metrics that appear frequently and deserve careful interpretation.<\/p>\n<p>Table 2: Classification metrics, definitions, and considerations\/pitfalls<\/p>\n<p>Metric<br \/>\nDefinitions<br \/>\nConsiderations\/Pitfalls<\/p>\n<p>Accuracy<br \/>\nProportion of all predictions that are correct<br \/>\nCan appear high when the model predicts the most common class (e.g., N2 sleep). Misleading in imbalanced datasets and should never be interpreted alone.<\/p>\n<p>Sensitivity \/ Recall (R)<br \/>\nTrue positive rate; proportion of actual cases correctly identified<br \/>\nCritical for patient safety when missing disease has high consequences (e.g., OSA). A model can have high accuracy but dangerously low recall.<\/p>\n<p>Precision (P) \/ PPV<br \/>\nProportion of predicted positives that are truly positive<br \/>\nOptimizing precision alone can underdiagnose disease. Low precision increases false positives and strains patient and clinician resources.<\/p>\n<p>F1-score<br \/>\nHarmonic mean of precision and recall<br \/>\nSummary metric balancing false positives and false negatives. Useful as a secondary score but can mask poor recall in rare conditions.<\/p>\n<p>Matthews Correlation Coefficient (MCC)<br \/>\nBalanced measure using all four confusion matrix elements (TP, TN, FP, FN*)<br \/>\nStrong metric for imbalanced classes and event detection. Often more informative than accuracy but less familiar to clinicians.<\/p>\n<p>Cohen\u2019s kappa (\u03ba)<\/p>\n<p>(Po \u2212 Pe) \/ (1 \u2212 Pe)<\/p>\n<p>Agreement between model and human scoring beyond chance**<\/p>\n<p>Useful for comparing AI with expert PSG scoring. Reflects clinical agreement, not overall model optimization. Should be interpreted alongside other metrics.<\/p>\n<p>Area under the curve of a receiver operating characteristic (AUC ROC)<br \/>\nAbility to discriminate across decision thresholds<br \/>\nCan appear high while still missing clinically important cases. Must be interpreted with recall and class balance.<\/p>\n<p>Area under the precision-recall curve (AUC-PR)<br \/>\nPrecision-recall performance emphasizing the positive class<br \/>\nMore sensitive to rare events than AUC ROC. Drops sharply with poor recall or precision; better reflects performance in imbalanced sleep datasets.<\/p>\n<p>*TP = true positive, TN = true negative, FP = false positive, FN = false negative<\/p>\n<p>**Po = observed agreement = accuracy (for the same unit), Pe = expected agreement by chance based on label marginals<\/p>\n<p>Conclusion<\/p>\n<p>AI models will increasingly shape how sleep clinicians diagnose, prognosticate, and document care. AI model performance depends on the objectives and the data used to train it. Metrics must be interpreted in the context of clinical risk to avoid bias, overdiagnosis, and missed treatment. Healthy skepticism entails aligning the highlighted metric with the clinical stakes. Clinicians do not need to become data scientists, but they must ask practical questions:<\/p>\n<p>What problem is the AI model solving?<br \/>\nWhich errors are most harmful?<br \/>\nAre development and deployment metrics transparent?<\/p>\n<p>In sleep medicine, clinicians act as informed gatekeepers rather than passive adopters.<\/p>\n<p>###<\/p>\n<p>Further thoughts<\/p>\n<p>Future\u00a0priorities\u00a0in\u00a0sleep\u00a0medicine\u00a0should include\u00a0promoting\u00a0the use of\u00a0interpretable AI model systems,\u00a0maintaining\u00a0oversight\u00a0throughout\u00a0the\u00a0AI\u00a0model lifecycle, and\u00a0adopting standardized\u00a0reporting\u00a0checklists.\u00a0These practices support fairness, reduce bias, and improve trust in clinical AI models.\u00a0<\/p>\n<p>This article appeared in volume 11, issue\u00a02\u00a0of\u202f<a class=\"Hyperlink SCXW156945097 BCX8\" href=\"https:\/\/aasm.org\/membership\/montage\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Montage magazine<\/a>.\u202f\u00a0<\/p>\n<p>References<\/p>\n<p>Kocak B, Klontzas ME, Stanzione A, Meddeb A, Demircio\u011flu A, Bluethgen C, Bressem KK, Ugga L, Mercaldo N, D\u00edaz O, Cuocolo R. Evaluation metrics in medical imaging AI: fundamentals, pitfalls, misapplications, and recommendations. Eur J Radiol Artif Intell. 2025;3:100030. <a href=\"https:\/\/doi.org\/10.1016\/j.ejrai.2025.100030\" target=\"_blank\" rel=\"noopener nofollow\">https:\/\/doi.org\/10.1016\/j.ejrai.2025.100030<\/a><\/p>\n<p>Further Reading<\/p>\n<p>Bandyopadhyay A, Bae C, Cheng H, et al. Smart sleep: what to consider when adopting AI-enabled solutions in clinical practice of sleep medicine.\u202fJ Clin Sleep Med. 2023;19(10):1823-1833. <a href=\"https:\/\/doi.org\/10.5664\/jcsm.10702\" target=\"_blank\" rel=\"noopener nofollow\">https:\/\/doi.org\/10.5664\/jcsm.10702<\/a><br \/>\nBandyopadhyay A, Oks M, Sun H, et al. Strengths, weaknesses, opportunities, and threats of using AI-enabled technology in sleep medicine: a commentary.\u202fJ Clin Sleep Med. 2024;20(7):1183-1191. <a href=\"https:\/\/doi.org\/10.5664\/jcsm.11132\" target=\"_blank\" rel=\"noopener nofollow\">https:\/\/doi.org\/10.5664\/jcsm.11132<\/a><br \/>\nGoldstein CA, Berry RB, Kent DT, et al. Artificial intelligence in sleep medicine: background and implications for clinicians.\u202fJ Clin Sleep Med. 2020;16(4):609-618. <a href=\"https:\/\/doi.org\/10.5664\/jcsm.8388\" target=\"_blank\" rel=\"noopener nofollow\">https:\/\/doi.org\/10.5664\/jcsm.8388<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"By Daniel Rongo, MD; Scott Ryals, MD; Mattina A. Davenport, PhD; and Trung Le, PhD, on behalf of&hellip;\n","protected":false},"author":2,"featured_media":61523,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[35214,35217,35218,24,35225,35219,25,18206,35228,35229,2537,35221,35215,35227,4360,35216,35220,35226,35224,35212,35222,35211,35223,35213],"class_list":["post-61522","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-aasm","tag-accreditation","tag-accredited","tag-ai","tag-american-academy","tag-apnea","tag-artificial-intelligence","tag-association","tag-circadian-rhythym","tag-disorder","tag-doctor","tag-insomnia","tag-membership","tag-mslt","tag-sleep","tag-sleep-apnea","tag-sleep-association","tag-sleep-deprivation","tag-sleep-disorders","tag-sleep-doctor","tag-sleep-lab","tag-sleep-medicine","tag-sleep-society","tag-technologist"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/61522","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=61522"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/61522\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/61523"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=61522"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=61522"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=61522"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}