The federal government is now using ChatGPT to make enforcement decisions that can freeze Medicaid funding to hospitals, nonprofits, and state agencies across all 50 states — and as of today, it has not published the tool’s error rate, disclosed its methodology, or specified when funds will actually be cut.
The program is called AERO — Audit Enforcement and Risk Oversight — and HHS launched it on May 21. Legal experts convened a compliance webinar on the initiative today, June 25, as compliance attorneys advised healthcare organizations to file Freedom of Information Act requests before responding to any HHS correspondence related to the program. The legal community is mobilizing because the stakes are immediate: AERO findings can trigger payment withholding, cost disallowances, suspension of awards, and debarment proceedings — the most severe of which permanently bars an organization from receiving federal funds.
What makes AERO distinct from prior federal fraud enforcement is the technology at its center. Gustav Chiarello, HHS Assistant Secretary for Financial Resources and the official leading the initiative, confirmed to the Wall Street Journal that ChatGPT and other large language models are being used to ingest and analyze audit records — documents that have historically piled up at the Federal Audit Clearinghouse largely unread. The AI does not make final enforcement decisions itself; it surfaces flags for human reviewers. But the question that compliance attorneys and policy experts are now asking is whether an AI system with a known tendency to generate plausible-sounding falsehoods should serve as the primary discovery mechanism for enforcement actions that can defund programs serving some of the country’s most vulnerable populations.
How HHS Turned an Unread Pile of Audits Into an AI Enforcement Pipeline
Under the Single Audit Act of 1984 and its 1996 amendments — codified in OMB Uniform Guidance at 2 CFR Part 200 — any state, local government, nonprofit, or university that spends $1 million or more in federal funds annually must submit a standardized financial compliance audit each year. These Single Audits cover financial statements, internal control assessments, and compliance with federal program requirements. They are submitted to the Federal Audit Clearinghouse and are public records.
For decades, the filings accumulated largely without consequence. Chiarello told the Wall Street Journal that audit reports landed “with a thud” and nobody acted on them. HHS’s review of its backlog before launching AERO found that some deficiencies had gone unaddressed for three, four, or five or more years; hundreds of grantees had not submitted required audits at all, with some delinquent by more than two years.
AERO changes the enforcement posture by giving ChatGPT and other large language models access to at least five years of these already-public records across all 50 states simultaneously. The AI scans for four categories: chronic noncompliance, repeat deficiencies, material weaknesses in internal controls, and delinquent audit submissions. Because the source documents are publicly filed records — not medical claims, not personal health information — the pipeline avoids the HIPAA and data-localization complications that have stalled prior federal AI projects.
The scope extends well beyond Medicaid. AERO covers all HHS-funded programs: Head Start childcare, federal medical research grants, addiction treatment services, and any other program administered by a grantee spending at least $1 million in federal funds annually. Public hospital systems, academic medical centers, federally qualified health centers, and behavioral health providers are all within scope.
HHS sent letters to all 50 state governors and state treasurers in May putting them on notice. The administration separately threatened to withhold Medicaid funds from all 50 states if they fail to comply with federal anti-fraud statutes. The Trump administration has already withheld hundreds of millions of dollars from Minnesota and more than $1 billion from California in Medicaid-related enforcement actions this year.
What Chiarello Says the AI Is Chasing
Chiarello has stated that he believes HHS is losing between $100 billion and $200 billion annually to wasteful or fraudulent spending — a range wider than the annual GDP of many countries. That figure is not a confirmed measurement of known fraud; it is his estimate of the department’s financial exposure, offered to justify deploying AI at the scale of the entire federal health funding system.
The enforcement consequences for non-compliant organizations are graduated but potentially severe. HHS has outlined four actions it may pursue against entities that fail to resolve audit findings: temporary withholding of payments until corrective action is taken; disallowance of costs already incurred; suspension or termination of awards; and initiation of debarment proceedings, which can permanently bar an organization from receiving federal funding.
The 2025 enforcement baseline established before AERO launched gives a sense of the administration’s existing capacity: the Centers for Medicare and Medicaid Services suspended $5.7 billion in Medicare payments, denied 122,658 claims, revoked 5,586 billing privileges, and generated 372 fraud referrals worth $3.7 billion that year — all before the AI audit sweep began.
Where AI Error Rate Meets Federal Enforcement
The sharpest concern from legal experts is not that AERO is auditing. It is that an AI system with a documented failure mode is generating the findings that trigger enforcement.
Generative AI models are known to hallucinate — to produce outputs that are confidently stated but factually wrong. In high-stakes domains, the consequences are documented: Stanford researchers found that leading legal AI tools hallucinate on roughly one in six queries. A New York lawyer was sanctioned for submitting ChatGPT-generated citations to a federal court; the cases did not exist. In security applications, AI false positives can trigger incident responses for threats that never existed.
In AERO’s context, an AI-generated false positive is not an embarrassment — it is a potential enforcement action. An audit finding that mischaracterizes a corrected deficiency as still open, or that attributes a documentation error to a grantee that has already resolved it, could trigger payment withholding against organizations that run childcare programs, addiction treatment centers, and federally qualified health clinics. The communities those organizations serve would bear the consequences of an error the AI committed.
Compounding the risk is a structural feature of how Single Audits work: findings against a state Medicaid agency do not stay contained to that agency. They cascade downstream to hospital subrecipients — private hospitals, clinics, and nonprofits that receive federal pass-through funds from the state but had no involvement in the underlying audit deficiency. A false positive generated against a state agency could create enforcement exposure for dozens of downstream organizations that the AI was not examining.
Ravi Patel, a pharmacist and former innovation lead at the Pittsburgh School of Pharmacy, told Drug Topics that AI’s ability to detect longitudinal fraud patterns is genuine — but the same symmetry applies to its errors. “The benefits and challenges of AI can apply to fraud in equal measure,” Patel said, noting that a single erroneous finding carries the same weight in an automated enforcement system as a legitimate one.
HHS acknowledged to the Associated Press a significant data error in a separate, earlier AI-assisted Medicaid fraud investigation in New York. The administration has not disclosed what corrective action was taken or whether findings issued in that investigation were revisited.
Legal Experts: File a FOIA Request Before Responding
AERO was announced through a press release and letters to governors — not through notice-and-comment rulemaking under the Administrative Procedure Act. That means no hospital, nonprofit, state agency, or advocacy organization had an opportunity to review the AI’s methodology, challenge its assumptions, or comment on its application before the program went live.
That procedural gap is now the foundation of the emerging legal strategy. Attorneys at the National Law Review advise every hospital or grantee that receives AERO-related correspondence from HHS to file a Freedom of Information Act request before responding — specifically seeking the AI’s methodology, training data, validation studies, bias assessments, and any documentation of compliance with OMB Memorandum M-25-21, the current federal AI governance framework issued in April 2025. Under M-25-21, agencies deploying AI with significant consequences for individuals are required to implement minimum risk management practices — including pre-deployment testing and an accessible human review process — and must report their compliance with these requirements to the White House Office of Management and Budget by September 2026. HHS has not publicly disclosed whether AERO has been assessed under that framework.
An AI-generated audit finding that is factually incorrect could be challenged under the APA’s “arbitrary and capricious” standard in § 706(2)(A), which requires courts to set aside agency action that lacks a reasoned factual basis. Courts have begun grappling with whether AI-generated findings satisfy that standard when the AI’s decision logic is a black box — a question a 2025 Harvard Law Review analysis described as presenting “novel issues” for administrative review.
What the DOJ Added on May 27
AERO did not arrive in isolation. Six days after the program launched, on May 27, Assistant Attorney General Brett Shumate issued a memorandum to DOJ prosecutors directing them to accelerate their review and enforcement of False Claims Act cases involving fraud in federally funded benefit programs. The Shumate memo — which follows President Trump’s March 2026 executive order establishing a task force to eliminate fraud — sets a goal of completing DOJ review of new benefits-fraud qui tam actions within the 60-to-120-day window prescribed by the False Claims Act.
Taken together, AERO and the Shumate memo describe a coordinated enforcement architecture: the AI scans audits and flags patterns; the DOJ closes the loop on individual cases through False Claims Act litigation. The coordination raises the stakes for any organization that receives an AERO letter and does not respond promptly with documented corrective action.
What This Means If Your State Receives Federal Health Funds
If you pay state or federal taxes — and your state receives Medicaid funding, which all 50 do — AERO is actively scanning financial records tied to programs you fund. The practical implications vary by situation.
Organizations with clean, current Single Audit records face low immediate risk. AERO targets chronic and repeat noncompliance — entities that have ignored the same deficiencies for years. A grantee that has submitted audits on time, documented corrective actions, and resolved flagged findings is not AERO’s primary target.
Organizations with longstanding unresolved findings are in a materially different position. A human reviewer buried in paperwork was, historically, a manageable risk. An AI system that surfaces five-year-old unresolved findings in seconds is not. The calculus has changed.
Hospitals, nonprofits, and grantees that have received federal funds through state Medicaid agencies — even if their own audits are clean — face indirect exposure through the downstream cascade risk described above. If AERO flags the state agency, its subrecipients may face scrutiny regardless of their own compliance history.
Legal and compliance experts are recommending three immediate actions: conduct an internal review of at least five years of Single Audit history and document any corrective actions already taken; confirm with your auditor which prior findings have been formally closed; and engage grant compliance counsel before responding to any HHS correspondence related to AERO.
ChatGPT as Infrastructure: The Broader Precedent
HHS deployed this LLM workflow without a formal procurement process, without a vendor selection competition, and without the multi-year authorization timeline that typically governs federal technology deployments at this scale. The agency used a publicly available tool on already-public records, which meant no new statutory authority was required and no data-sharing restrictions applied.
That combination — commercial AI, public documents, existing legal authority — sets a procurement template that other federal agencies can replicate. The next, harder application is bringing AI into pre-payment claims integrity for actual Medicaid and Medicare billing data, where the source documents are not public and the protected health information complications are substantial. AERO demonstrates that the federal government can ship an LLM enforcement workflow faster than it can navigate a traditional technology acquisition — a capability with implications well beyond this program.
HHS has stated it does not yet have an official estimate for the financial savings AERO will produce. What is already clear is that the federal government has crossed a threshold. ChatGPT is not just a consumer product or an enterprise productivity tool. It is now part of the infrastructure through which the federal government determines where public health dollars go — and where they get taken back. Whether that infrastructure can produce findings accurate enough to survive the legal scrutiny it will inevitably face is a question the litigation pipeline is beginning to answer.
Frequently Asked Questions
What is the HHS AERO program and which organizations does it affect?
AERO — the Audit Enforcement and Risk Oversight initiative — is a program HHS launched on May 21, 2026, that uses ChatGPT and other large language models to scan at least five years of Single Audit Act compliance records across all 50 states. Any organization that receives $1 million or more in federal funds annually and is required to file an annual Single Audit is within scope. That includes state Medicaid agencies, nonprofit healthcare providers, public hospital systems, federally qualified health centers, academic medical centers, Head Start programs, addiction treatment providers, and federal research grant recipients. Enforcement consequences range from temporary payment holds to permanent debarment from federal programs.
How is ChatGPT being used to detect Medicaid fraud?
ChatGPT and related large language models are ingesting Single Audit reports — publicly filed documents available in the Federal Audit Clearinghouse — to flag patterns of chronic noncompliance, repeat internal control deficiencies, material weaknesses, and delinquent audit submissions. The AI surfaces findings for human HHS reviewers, who then determine which states and grantees to pursue. The source documents contain no personal health information; the pipeline operates on financial compliance filings, not on medical claims or patient records. The AI identifies patterns across years of filings simultaneously — something that was practically impossible for human auditors given the volume of documents involved.
What happens if a state or grantee fails an HHS AERO audit review?
HHS has outlined a graduated set of enforcement actions: temporary withholding of federal payments until the deficiency is corrected; disallowance of costs already incurred; suspension or termination of federal awards; and initiation of debarment proceedings, which can permanently bar an organization from receiving federal funding. The letters HHS sent to all 50 governors in May did not specify a deadline for corrective action or a threshold at which funds would be withheld. Legal experts recommend that any organization receiving AERO-related correspondence from HHS file a Freedom of Information Act request for the AI’s methodology before responding, and engage grant compliance counsel.
Can AI hallucinations or errors result in wrongful enforcement actions under AERO?
Yes, and this is the central legal concern experts are raising. Generative AI models have a documented tendency to produce factually incorrect outputs with apparent confidence — a failure mode called hallucination. In AERO’s context, a false positive could trigger enforcement action against an organization that has already corrected the deficiency the AI flagged, or against a hospital or nonprofit that receives state-passed-through federal funds and had no involvement in the deficiency AERO identified at the state agency level. Legal analysts have noted that an AI-generated enforcement finding that is factually incorrect could be challenged under the Administrative Procedure Act’s “arbitrary and capricious” standard, which requires agency actions to have a reasoned factual basis. HHS has not disclosed the AI’s error rate, published validation studies, or established a public appeals process for organizations that dispute AERO findings.