
ChatGPT application displayed on a smartphone screen, highlighting its growing role in everyday life and the increasing reliance on artificial intelligence tools that are shaping decision making and automating tasks, amid rising concerns over AI monopolizing certain jobs and prompting companies to reduce their workforce, in Tunis,Tunisia on May 5,2026.
Imen Ben Youssef/Getty Images
The federal government is now using ChatGPT to scan five years of compliance records for every hospital, nonprofit, university, and state Medicaid agency that receives federal health dollars — and it has not disclosed how often the tool gets it wrong. The program is called AERO, the Audit Enforcement and Risk Oversight initiative, and the Department of Health and Human Services launched it on May 21, 2026, without a public solicitation, without a published validation study, and without specifying when enforcement will begin. What HHS has specified is the consequence of being flagged: organizations can lose their federal funding entirely — temporarily through payment holds, or permanently through debarment proceedings that bar them from all federal programs.
The architectural problem with this deployment is not merely that an error rate is missing from a disclosure form. Large language models like ChatGPT generate plausible-sounding outputs by predicting text patterns; they have no internal mechanism to signal when they are uncertain about a specific factual question. A study published in Nature in April 2026 formally established that hallucination is statistically inevitable in pretrained models for non-learnable facts — precisely the category of question AERO poses: did this grantee resolve this deficiency, or did they not? An audit compliance status in a federal database looks textually similar whether the corrective action was filed six months later and not captured in the scan, or never filed at all. ECRI, the independent patient safety organization, named ChatGPT misuse as the single greatest health technology hazard of 2026 in its annual Top 10 Health Technology Hazards report, noting explicitly that general-purpose tools “are not regulated as medical devices and have not been validated for healthcare purposes.” HHS has deployed one for enforcement decisions that can permanently end an organization’s federal funding.
A critical clock is now running. Under OMB Memorandum M-25-21, federal agencies deploying “High Impact AI” — defined as AI with significant consequences for individuals’ access to government benefits or services — must report their minimum risk management practices to the White House Office of Management and Budget by September 22, 2026. Those practices include pre-deployment testing, continuous monitoring, and an accessible human review and remediation process. Any High Impact AI use case that is not compliant with those requirements must be discontinued. HHS has not publicly disclosed whether AERO has been assessed under that framework, and whether it has a compliant human review process in place. If the September 22 report confirms it does not, HHS may be legally required to shut the program down.
What AERO Actually Does — and What It Does Not
The program’s mechanics are more targeted than public perception suggests. AERO does not directly scan Medicare or Medicaid claims for fraudulent billing. It reads Single Audit reports: the annual compliance filings that any state, local government, nonprofit, or higher education institution must submit to the Federal Audit Clearinghouse if it spends $1 million or more in federal funds annually. These documents, publicly searchable on the GSA Federal Audit Clearinghouse, have historically gone largely unreviewed. Gustav Chiarello, the HHS Assistant Secretary for Financial Resources and Chief Financial Officer who is leading the program, described the problem candidly: audits “land with a thud and no one does anything about it,” according to reporting by the Wall Street Journal. HHS’s review of its backlog before launching AERO found some deficiencies had gone unaddressed for five or more years; hundreds of grantees had not submitted required audits at all, some delinquent by more than two years.
AERO’s LLM pipeline ingests these already-public compliance documents and scans for four categories: chronic noncompliance, repeat deficiencies, material weaknesses in internal controls, and delinquent audit submissions. Because the source documents are publicly filed financial compliance reports — not medical claims and not personal health information — the deployment sidesteps the HIPAA and data-localization complications that have stalled prior federal AI projects. This is the procurement innovation at AERO’s core: HHS built a consequential enforcement pipeline from off-the-shelf ChatGPT applied to data it already held, without a public request for proposals, without a multi-year authorization process, and in a matter of months.
Chiarello told the Wall Street Journal the initiative grew out of his review of childcare fraud in Minnesota. He estimated HHS has between $100 billion and $200 billion in annual wasteful or fraudulent spending — a figure he offered not as a measured amount of confirmed fraud, but as an estimate of the department’s exposure, offered to justify the scale of the AI deployment. That distinction matters for readers evaluating how much weight to give AERO’s early findings.
Who Is Inside AERO’s Reach
The scope of organizations inside AERO’s reach is broader than most Americans realize. Any entity spending $1 million or more in federal funds annually falls under the Single Audit Act. That includes state Medicaid agencies, public hospital systems, academic medical centers, federally qualified health centers, Head Start programs, addiction treatment providers, and federal research grant recipients. A rural nonprofit running a substance abuse clinic, a state children’s health agency, and a public university with federal research grants are all reading from the same list of potential enforcement targets.
The consequences are graduated but potentially severe. HHS has outlined a spectrum of enforcement actions: temporarily withholding payments until corrective action is taken, disallowing specific costs already incurred, suspending or terminating program awards, and — at the far end — initiating debarment proceedings that would permanently bar an organization from receiving federal funds. HHS sent letters to all 50 state governors and treasurers putting them on notice of the new initiative. Critically, those letters did not include a deadline for corrective action or a timeline for when HHS might begin withholding funds.
A legal wrinkle documented by the National Law Review makes the exposure even more unpredictable: AERO findings against a state Medicaid agency can flow downstream to hospital subrecipients that had no involvement in the underlying audit deficiency. An organization in full compliance today may face consequences from findings about an intermediary agency over which it has no control.
Does the Federal Government’s Own AI Playbook Allow This?
The central legal challenge to AERO is not primarily about whether ChatGPT makes mistakes. It is about whether the process HHS used to deploy it is lawful, and whether the outputs it generates can legally ground enforcement actions.
AERO was announced through a press release and letters to governors — not through the notice-and-comment rulemaking process that the Administrative Procedure Act requires for significant changes in enforcement policy. No organization had the opportunity to review the AI methodology, challenge its assumptions, or comment before it was applied to years of historical compliance data. Legal scholars at the Yale Journal on Regulation and the Regulatory Review have concluded that agencies “will not be able to rely solely on today’s most ubiquitous forms of AI — namely, those based on ChatGPT and similar large language models — to avoid their obligation under the APA’s arbitrary and capricious standard,” which requires agencies to produce a reasoned, documented factual basis for their decisions. An opaque AI output that cannot explain its own reasoning does not satisfy that standard, according to analyses from multiple legal scholars including Cary Coglianese of the University of Pennsylvania.
The National Law Review also noted that HHS’s own published internal standards — its Trustworthy AI Playbook — require bias testing, human oversight, transparency, and OMB pre-clearance before deploying rights-impacting AI. There is no public evidence that AERO met any of those requirements before it launched.
The Trump administration acknowledged to the Associated Press that it used incorrect data in at least one case involving a New York Medicaid fraud investigation — a concrete example of data errors translating into real enforcement action. Rob Weissman, co-president of the consumer advocacy group Public Citizen, raised broader questions about whether the administration is applying AERO’s findings evenhandedly across states of different political affiliations. Before AERO launched, the administration had withheld hundreds of millions of dollars in Medicaid funds from Minnesota and more than $1 billion from California; critics have argued enforcement has disproportionately targeted states with Democratic governors.
How Does ChatGPT Actually Detect Noncompliance — and Where Can It Fail?
The pipeline that ChatGPT runs in AERO is, technically, a document pattern-recognition task, not a fraud determination. The Federal Audit Clearinghouse stores Single Audit reports, including each organization’s Data Collection Form (Form SF-SAC), which summarizes the audit findings in standardized fields. The LLM scans these structured and semi-structured records to identify patterns — chronic repeat findings across multiple years, material weakness designations, missing submissions — that would be nearly impossible for human auditors to catch across tens of thousands of filings simultaneously.
That is a legitimate and well-suited task for a large language model. The problem emerges at the margin cases. When an organization corrected a compliance deficiency and submitted documentation that was filed in a subsequent period not captured in AERO’s five-year scan window, the model may read the earlier finding as unresolved. When a prior-period finding was addressed through a corrective action plan but the FAC record does not clearly mark it as closed — a problem GAO documented in a 2024 report on FAC data quality, finding that the clearinghouse “can’t identify recipients that are required to submit a single audit but didn’t” and has known completeness gaps — the model has no mechanism to flag its own uncertainty. It will produce a confident-sounding output that the deficiency is unresolved, even when the factual record is ambiguous.
This is not a hypothetical failure mode. The Nature paper on LLM hallucination established that the failure is most likely to occur precisely where the factual record is inconsistent, incomplete, or spread across multiple document versions — which describes a significant portion of the Federal Audit Clearinghouse’s historical holdings. The problem for organizations facing an AERO enforcement letter is that the AI’s output looks like a definitive finding, not a probability estimate. Without a published error rate or a disclosed methodology, there is no public basis to evaluate how often AERO gets the borderline cases wrong.
What the OpenAI-Government Relationship Means for AERO’s Oversight
HHS’s reliance on ChatGPT for AERO arrives at a commercially and legally complicated moment for OpenAI. The company confidentially filed draft IPO registration documents with the SEC on May 22, 2026, with Goldman Sachs and Morgan Stanley as lead underwriters. As of July 2026, advisers are weighing whether to target a 2027 listing rather than the originally expected fourth quarter of 2026, with CEO Sam Altman stating he views any valuation below $1 trillion as unacceptable. The company has also entered early discussions about giving the U.S. government a 5% equity stake — a proposal reported by the Financial Times on July 2, 2026 that would require Congressional approval and remains in a conceptual stage.
A coalition of state attorneys general subpoenaed OpenAI on June 13, 2026, demanding documents on topics including the company’s handling of consumer data and health data, and model behavior — a probe that arrived the same week OpenAI confirmed its confidential SEC filing. The multi-state investigation adds regulatory complexity to a company simultaneously seeking a landmark public offering and deepening its federal government relationships.
The structural concern is one that legal and governance scholars have not yet resolved: when the government is a major customer of an AI platform and is simultaneously under discussion as a prospective equity holder in that platform, the oversight mechanisms that typically govern federal technology deployments — independent audit, procurement competition, validation requirements — lose some of their independence. AERO illustrates what that looks like in practice: a major federal enforcement program built on a commercial AI product, deployed without an RFP, without a published validation study, and with enforcement consequences for thousands of organizations, by an agency that is a major customer of the platform whose accuracy it has not assessed.
AERO as a Government AI Procurement Template
One of the underreported aspects of AERO is the precedent it sets for how federal agencies buy and deploy AI. HHS built an enforcement pipeline from off-the-shelf ChatGPT applied to data it already held, without a public request for proposals and without a multi-year authorization process. Chiarello told reporters that other federal departments could easily replicate the model: “It would be fairly easy for the other agencies to use our technology and jump on it.” If the template scales, AERO would mark not just a program launch but a new procurement paradigm for AI in government — one in which agencies deploy commercial AI against public data they already possess, bypassing the standard federal acquisition stack entirely.
That precedent is what makes the September 22, 2026 M-25-21 deadline consequential beyond AERO itself. If HHS’s September 22 report to OMB discloses that AERO lacks a human review process, lacks pre-deployment validation, and lacks a documented error rate — and if OMB acts on M-25-21’s requirement that non-compliant High Impact AI systems be discontinued — the outcome would establish that agencies cannot simply declare compliance-by-silence and deploy consequential AI. If, instead, the report finds AERO meets the requirements or receives a waiver, it will establish a precedent for how little process a federal agency needs before deploying an LLM with enforcement consequences.
Legal experts expect the first False Claims Act cases connected to AERO findings to move through DOJ in the coming months. The DOJ issued its own enforcement acceleration memo on May 27, 2026 — six days after AERO launched — directing prosecutors to target benefits-fraud cases for completion within 60 to 120 days. How courts treat AI-generated compliance flags in those cases — as sufficient evidentiary basis, or as black-box outputs that cannot satisfy the APA’s reasoned-basis standard — will set precedent for every subsequent federal AI enforcement program.
The audit landed with a thud for years. Now it has the attention of a large language model — one whose error rate has not been published, whose human review process has not been specified, and whose legal standing in federal enforcement proceedings has not been tested.
Frequently Asked QuestionsWhat does AERO actually scan, and can it directly freeze my organization’s federal funding?
AERO reads Single Audit reports — publicly filed annual compliance documents submitted to the Federal Audit Clearinghouse by any organization spending $1 million or more in federal funds. It does not scan Medicaid or Medicare claims directly. However, what it finds can trigger enforcement: HHS can temporarily withhold payments, disallow costs, suspend or terminate awards, or initiate debarment proceedings. Debarment — permanent exclusion from all federal programs — is the most severe outcome. A letter from HHS tied to AERO does not represent a final enforcement decision, but compliance attorneys uniformly advise filing a Freedom of Information Act request seeking the AI’s methodology, validation studies, and OMB clearance documentation before responding to any such correspondence.
Why is it a problem that ChatGPT makes enforcement decisions without a published error rate?
Large language models are architecturally incapable of knowing when they don’t know something. They generate the most statistically plausible output based on text patterns, without any internal signal when a specific factual judgment — did this grantee resolve this audit finding or not? — is genuinely uncertain. A GAO 2024 report documented that the Federal Audit Clearinghouse itself has completeness gaps and data quality issues, meaning the underlying records AERO is scanning have known inaccuracies. When the AI processes incomplete or ambiguous compliance records and produces a confident-sounding finding of noncompliance, there is currently no disclosed process for the organization to challenge that finding before enforcement begins. Without a published error rate or a validated methodology, organizations cannot evaluate how often that kind of confident error occurs.
What is the September 22, 2026 deadline and what happens if AERO fails it?
Under OMB Memorandum M-25-21, federal agencies deploying “High Impact AI” — AI that has significant consequences for access to government benefits — must report their minimum risk management practices to OMB by September 22, 2026. Those practices include pre-deployment testing, continuous monitoring, and an accessible human review process. If an agency’s use case does not comply with those minimum practices, M-25-21 requires it to be discontinued. AERO almost certainly qualifies as High Impact AI, since its findings can result in organizations losing federal funding. If HHS’s September 22 report to OMB discloses that AERO lacks these safeguards, it would create either a legal obligation to shut the program down or a governance confrontation between OMB and HHS. That report is the most concrete near-term test of whether the federal government’s own AI rules apply to its own enforcement tools.
Can an organization challenge an AERO finding in court?
Not easily, and not yet — but legal scholars expect the mechanisms to emerge from the first wave of False Claims Act cases connected to AERO findings. The primary legal theory is the Administrative Procedure Act’s “arbitrary and capricious” standard: courts must set aside agency action that lacks a reasoned factual basis, and legal scholars including Cary Coglianese of Penn and analysts at the National Law Review argue that a black-box AI output without a disclosed methodology cannot satisfy that standard. The first step — before any lawsuit — is the FOIA request: filing for AERO’s AI methodology, training data, validation studies, and OMB pre-clearance documents establishes an administrative record, surfaces methodological gaps, and positions any challenge on the strongest available factual footing. Courts in analogous AI enforcement challenges have used FOIA-obtained records as evidentiary foundations for invalidating agency actions.