{"id":51014,"date":"2026-05-26T04:40:12","date_gmt":"2026-05-26T04:40:12","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/51014\/"},"modified":"2026-05-26T04:40:12","modified_gmt":"2026-05-26T04:40:12","slug":"data-privacy-and-ai-progress","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/51014\/","title":{"rendered":"Data Privacy and AI Progress"},"content":{"rendered":"<p>To improve AI performance, regulators must loosen restrictions on data sharing.<\/p>\n<p>Every law reflects a cost benefit analysis made at the time of enactment. With respect to many privacy laws, that analysis is out of date. The costs of collecting, <a href=\"https:\/\/ourworldindata.org\/grapher\/historical-cost-of-computer-memory-and-storage\" rel=\"nofollow noopener\" target=\"_blank\">storing<\/a>, and analyzing data in sensitive contexts, such as in health care and educational settings, once vastly <a href=\"https:\/\/drexel.edu\/law\/lawreview\/issues\/Archives\/v8-2\/zeide\/?utm_source=chatgpt.com\" rel=\"nofollow noopener\" target=\"_blank\">outweighed<\/a> the benefits. In the 1970s, for example, it likely did not <a href=\"https:\/\/ourworldindata.org\/grapher\/historical-cost-of-computer-memory-and-storage\" rel=\"nofollow noopener\" target=\"_blank\">make sense<\/a> to collect any more information than was necessary to address a patient\u2019s immediate needs, let alone to <a href=\"https:\/\/ourworldindata.org\/grapher\/historical-cost-of-computer-memory-and-storage\" rel=\"nofollow noopener\" target=\"_blank\">store<\/a> that information for very long, due to high costs. The risks of that information being stolen or leaked were much graver than any positive outcomes. The same is not true today.<\/p>\n<p>Advances in artificial intelligence (AI) have <a href=\"https:\/\/arxiv.org\/pdf\/2001.08361\" rel=\"nofollow noopener\" target=\"_blank\">changed<\/a> the math. It is now the case that sharing is caring\u2014caring for your future self as well as caring for the wellbeing of others. For example, data <a href=\"https:\/\/pmc.ncbi.nlm.nih.gov\/articles\/PMC8827025\/\" rel=\"nofollow noopener\" target=\"_blank\">collected<\/a> during a patient\u2019s youth may drastically <a href=\"https:\/\/pmc.ncbi.nlm.nih.gov\/articles\/PMC8827025\/\" rel=\"nofollow noopener\" target=\"_blank\">improve<\/a> diagnoses in adulthood. Moreover, that data may also <a href=\"https:\/\/academic.oup.com\/jamia\/article\/30\/12\/2072\/7259105?\" rel=\"nofollow noopener\" target=\"_blank\">train<\/a> and improve models that result in better treatment for others. Those gains, however, will not occur if people continue to see data as something to hoard rather than to share. Similarly, if outdated laws <a href=\"https:\/\/scholarship.law.ufl.edu\/facultypub\/1274\/?\" rel=\"nofollow noopener\" target=\"_blank\">remain<\/a> on the books, then AI progress will be delayed, leading to lives being lost and educational gains going unrealized.<\/p>\n<p>Of course, there are still costs to permitting more liberal data collection in sensitive domains. But those costs must be put in the context of a health care system in which medical errors <a href=\"https:\/\/pubmed.ncbi.nlm.nih.gov\/27143499\/\" rel=\"nofollow noopener\" target=\"_blank\">run<\/a> rampant and an education system in which many kids slip through the cracks. AI will not be able to address those shortcomings absent legal reforms and cultural shifts in how we think about data sharing.<\/p>\n<p>That is because data are the basis on which AI systems learn what counts as a pattern and what counts as noise. The modern literature on machine learning has <a href=\"https:\/\/arxiv.org\/pdf\/2001.08361\" rel=\"nofollow noopener\" target=\"_blank\">shown<\/a>, with striking consistency, that model performance improves as training data grows and that even very large models <a href=\"https:\/\/arxiv.org\/pdf\/2203.15556\" rel=\"nofollow noopener\" target=\"_blank\">underperform<\/a> when they are trained on too few data. Quantity is only part of the story, however. A system <a href=\"https:\/\/proceedings.mlr.press\/v162\/bansal22b\/bansal22b.pdf?utm\" rel=\"nofollow noopener\" target=\"_blank\">trained<\/a> on narrow or unrepresentative data may appear accurate in familiar settings but fail when it encounters the harder cases that matter most in practice: unusual presentations, minority populations, atypical learning profiles, or circumstances that differ from the environment in which the model was developed.<\/p>\n<p>Rich, diverse, and well-curated datasets <a href=\"https:\/\/proceedings.mlr.press\/v162\/bansal22b\/bansal22b.pdf?utm\" rel=\"nofollow noopener\" target=\"_blank\">make<\/a> models more capable and more dependable by reducing the odds that a system has merely learned a shortcut that collapses outside the lab. In sensitive domains, then, building AI that is reliable enough to deserve public trust <a href=\"https:\/\/proceedings.mlr.press\/v162\/bansal22b\/bansal22b.pdf?utm\" rel=\"nofollow noopener\" target=\"_blank\">requires<\/a> data.<\/p>\n<p>Education law <a href=\"https:\/\/drexel.edu\/law\/lawreview\/issues\/Archives\/v8-2\/zeide\/?utm_source=chatgpt.com\" rel=\"nofollow noopener\" target=\"_blank\">offers<\/a> a concrete example of a data-limiting regime built on the assumption that restraint equals safety. California\u2019s <a href=\"https:\/\/leginfo.legislature.ca.gov\/faces\/codes_displayText.xhtml?lawCode=BPC&amp;division=8.&amp;title=&amp;part=&amp;chapter=22.2.&amp;article\" rel=\"nofollow noopener\" target=\"_blank\">Student Online Personal Information Protection Act<\/a> bars operators of K-12 online services from using covered information to amass a profile about a student except in furtherance of K-12 school purposes, forbids selling student information, and restricts disclosure. The statute <a href=\"https:\/\/leginfo.legislature.ca.gov\/faces\/codes_displayText.xhtml?lawCode=BPC&amp;division=8.&amp;title=&amp;part=&amp;chapter=22.2.&amp;article\" rel=\"nofollow noopener\" target=\"_blank\">preserves<\/a> room for some socially valuable uses. For example, it <a href=\"https:\/\/leginfo.legislature.ca.gov\/faces\/codes_displayText.xhtml?lawCode=BPC&amp;division=8.&amp;title=&amp;part=&amp;chapter=22.2.&amp;article\" rel=\"nofollow noopener\" target=\"_blank\">allows<\/a> educators to use deidentified information to improve educational products and permits the use of pupil data for adaptive or customized learning. But the architecture of the law still <a href=\"https:\/\/drexel.edu\/law\/lawreview\/issues\/Archives\/v8-2\/zeide\/?utm_source=chatgpt.com\" rel=\"nofollow noopener\" target=\"_blank\">reflects<\/a> an older intuition that student data should remain tethered to its immediate instructional purpose rather than contribute to broader systems of learning over time. That instinct is understandable in a world worried about commercialization and misuse. In a world of data-hungry AI tools, however, limitations of this sort can <a href=\"https:\/\/d-nb.info\/1176371517\/34?\" rel=\"nofollow noopener\" target=\"_blank\">make<\/a> it harder to build systems that learn from longitudinal and cross-context patterns. These are the very patterns that may be <a href=\"https:\/\/www.cse.unr.edu\/~sjose\/Papers%20Referenced\/Student%20modeling%20approaches%20A%20literature%20review%20for%20the%20last%20decade.pdf?utm_source=chatgpt.com\" rel=\"nofollow noopener\" target=\"_blank\">needed to<\/a> identify effective interventions, detect struggling students earlier, and develop tutors that work well for more than the easiest cases.<\/p>\n<p>Health care law offers the same pattern. The <a href=\"https:\/\/scholarship.law.ufl.edu\/facultypub\/1274\/?\" rel=\"nofollow noopener\" target=\"_blank\">Privacy Rule<\/a> in the <a href=\"https:\/\/www.govinfo.gov\/content\/pkg\/PLAW-104publ191\/pdf\/PLAW-104publ191.pdf\" rel=\"nofollow noopener\" target=\"_blank\">Health Insurance Portability and Accountability Act<\/a> (HIPAA) is built around bounded use, not maximal learning. Outside of treatment, covered entities generally must limit their uses and disclosures of protected health information to specified categories, and the rule\u2019s \u201cminimum necessary\u201d standard <a href=\"https:\/\/www.ecfr.gov\/current\/title-45\/subtitle-A\/subchapter-C\/part-164\/subpart-E\/section-164.502\" rel=\"nofollow noopener\" target=\"_blank\">requires<\/a> reasonable efforts to limit data to what the immediate purpose demands. Research can proceed, but broader reuse often <a href=\"https:\/\/www.hhs.gov\/hipaa\/for-professionals\/security\/laws-regulations\/index.html\" rel=\"nofollow noopener\" target=\"_blank\">depends<\/a> on patient authorization, a waiver by an institutional review board or privacy board, or deidentification. That design made sense when the law\u2019s central concern was disclosure. It makes less sense when accurate medical AI <a href=\"https:\/\/academic.oup.com\/jamia\/article\/30\/12\/2072\/7259105?\" rel=\"nofollow noopener\" target=\"_blank\">depends<\/a> on large, longitudinal, and clinically diverse datasets. These datasets would let a model learn from rare presentations, delayed complications, and patients whose care spans multiple years and providers.<\/p>\n<p>Excess data hoarding is now an anti-social behavior. Imagine patients who refuse to allow a doctor to use AI to transcribe their appointment. Not only would those patients miss out on the benefits of a more accurate summary of their meeting, but they would also <a href=\"https:\/\/pubmed.ncbi.nlm.nih.gov\/39688515\/\" rel=\"nofollow noopener\" target=\"_blank\">cause<\/a> other patients to have less time with the doctor, who will now have to spend time putting pen to paper rather than sitting down with patients. The same is true in education: When schools <a href=\"https:\/\/www.cse.unr.edu\/~sjose\/Papers%20Referenced\/Student%20modeling%20approaches%20A%20literature%20review%20for%20the%20last%20decade.pdf?utm_source=chatgpt.com\" rel=\"nofollow noopener\" target=\"_blank\">wall off<\/a> the diverse performance and interaction data that tutoring systems use to model student knowledge, they make it harder to build AI tutors that can adapt to different paces, misconceptions, and learning styles rather than merely serving the median student.<\/p>\n<p>None of this is an argument for treating sensitive data carelessly. The case for broader use is strongest only if it is <a href=\"https:\/\/www.ncbi.nlm.nih.gov\/books\/NBK613199\/?utm_source=chatgpt.com\" rel=\"nofollow noopener\" target=\"_blank\">matched<\/a> with a much higher bar for cybersecurity and a real respect for patient autonomy. HIPAA already <a href=\"https:\/\/www.hhs.gov\/hipaa\/for-professionals\/security\/laws-regulations\/index.html\" rel=\"nofollow noopener\" target=\"_blank\">emphasizes<\/a> cybersecurity by requiring administrative, physical, and technical safeguards for electronic protected health information. Any reform worth enacting should strengthen those protections rather than relax them. Patients should also <a href=\"https:\/\/www.ncbi.nlm.nih.gov\/books\/NBK613199\/?utm_source=chatgpt.com\" rel=\"nofollow noopener\" target=\"_blank\">retain<\/a> meaningful control: clear notice, genuine choices where feasible, and firm consequences for misuse.<\/p>\n<p>But that is not an argument for reflexive non-sharing. Other sectors already operate on the premise that people may entrust highly sensitive information to regulated institutions when the social gains are large. Consumers <a href=\"https:\/\/www.ftc.gov\/business-guidance\/resources\/ftc-safeguards-rule-what-your-business-needs-know\" rel=\"nofollow noopener\" target=\"_blank\">do<\/a> so in finance, where federal law requires covered institutions to maintain comprehensive information-security programs, because secure data flows make modern payments, credit, and fraud detection possible. Aviation <a href=\"https:\/\/asrs.arc.nasa.gov\/overview\/summary.html\" rel=\"nofollow noopener\" target=\"_blank\">does<\/a> the same through confidential safety data sharing programs that give regulators and operators a fuller picture of systemic risk. Health care should pursue the same balance: strong security, real agency, and broad enough data sharing to build tools that are worth trusting.<\/p>\n<p>The question, then, is not whether sensitive data should remain protected. It is whether the law will continue to protect it in ways that reflect yesterday\u2019s tradeoffs rather than today\u2019s realities. Rules built for an analog world treated restraint as the safest course because the gains from broader use were speculative, and the risks of misuse were obvious. AI has altered that balance. In both health care and education, the refusal to collect, retain, and responsibly share data now carries its own costs: worse diagnoses, weaker interventions, less personalized instruction, and slower improvement for everyone who comes next. The better path is not carelessness with data but a shift from data minimization to data stewardship. An effective policy would pair strong cybersecurity, real respect for individual choice, and serious penalties for abuse with a legal framework that permits socially valuable learning on a meaningful scale. If the law keeps treating sensitive data as something to be locked away, it will safeguard privacy while entrenching mistakes and inequality.<\/p>\n<p> <img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/05\/FrazierHeadshot.png\" alt=\"Kevin T. Frazier \" class=\"photo\" height=\"80\" width=\"80\"\/> <img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/05\/gloverheadshot.jpg\" alt=\"Gillian Glover\" class=\"photo\" height=\"80\" width=\"80\"\/><\/p>\n","protected":false},"excerpt":{"rendered":"To improve AI performance, regulators must loosen restrictions on data sharing. Every law reflects a cost benefit analysis&hellip;\n","protected":false},"author":2,"featured_media":51015,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,25,4060,507,76,3165],"class_list":["post-51014","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-artificial-intelligence","tag-artificial-intelligence-regulation","tag-data-privacy","tag-education","tag-hipaa"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/51014","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=51014"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/51014\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/51015"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=51014"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=51014"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=51014"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}