Claude Science

Claude Science
anthropic.com

Anthropic acknowledged today that its billing system generated phantom invoices of $1.67 million and then $16.6 million against a South Korean developer who was on the free tier, had never spent a dollar on the API, and had no payment card registered to the account — a billing error the company subsequently confirmed along with its statement that no money was collected. The confirmation came not before repeated charge attempts blocked the developer’s primary credit card and left him spending four days and roughly 18 support emails trying to get written confirmation the invoices were void. That incident, confirmed by Anthropic on July 12, arrives in the middle of a broader and older story: an audited enterprise overcharge pattern, a proposed class action, and a structural billing architecture that makes it nearly impossible for any customer to independently verify what they actually owe.

The developer, who posts online as remy_notes, described discovering the first invoice — a “payment unsuccessful” notice for $1,669,875.90 — on the night of July 7 and initially assuming it was a phishing attempt. Both emails passed every authenticity check he ran: the sender domain was Anthropic’s, and the Stripe payment link resolved to Anthropic’s official billing infrastructure. Screenshots he shared on Threads show his KB Kookmin check card declining two overseas charge attempts on July 8 — one at 6:01 PM and one at 7:41 PM — from a US merchant listed as “ANTHROP,” each rejected not because the bank detected fraud but because the sums exceeded the card’s per-transaction limit. The invoices grew roughly tenfold in 24 hours. Anthropic told him the incident “was not the result of unauthorized access” but has not disclosed the root cause of the billing system error.

Auditors Found $1.7M in Overcharges Before This Incident Broke

The phantom invoice is extreme, but the underlying problem it represents was already documented before it happened. Between March and June of this year, AI billing audit startup Vaudit reviewed $34 million in AI invoices submitted by 60 enterprise customers and found approximately $1.7 million in mistaken overcharges — a billing error rate of roughly five percent. The majority of the discrepancies were tied to Anthropic’s Claude Code product, though OpenAI services also appeared in the findings. Vaudit’s client roster for this audit included Panasonic, HP, and Honda.

Vaudit CEO Michael Hahn, a former Oracle director, described recurring billing failure patterns in The Information’s audit report: customers charged at premium rates for older, cheaper models they were actually using; AI agents and chatbots that failed to complete requests or returned error messages but were still billed; and what the company calls “retry storms” — autonomous agents that repeatedly retry failed tasks in the background, each attempt generating a new charge the customer never authorized. After Vaudit and its clients formally challenged the invoices, providers — including Amazon, Google, Microsoft, Anthropic, and OpenAI — refunded approximately 80 percent of the disputed amounts, typically within 48 to 96 hours. Hahn described the companies as “incredibly cooperative” once clear-cut errors were presented.

That 80 percent recovery rate sounds reassuring until you ask the inverse question: how many customers without an auditor never noticed and never recovered the remaining 20 percent?

Anthropic, for its part, pushed back on the framing. A company spokesperson said Anthropic does not charge for incomplete requests or error messages, does not route users to higher-priced models without their knowledge, and called systematic overbilling “not a widespread problem.” On the phantom invoice: “No money left your account. Our payment processor attempted a charge at the invalid amount and it was declined. Nothing was collected, and you owe nothing.”

Token Billing Is Structurally Unverifiable — That Is the Core Problem

Understanding why AI billing errors are so hard to catch requires understanding how the billing works. Every Claude request is priced in tokens — the internal units a language model processes for each prompt and response. Input tokens (the text sent to the model) and output tokens (the model’s response) are counted separately and multiplied by per-model rates. Claude Fable 5, Anthropic’s most capable publicly available model, carries rates of $10 per million input tokens and $50 per million output tokens — the highest the company has ever published for a generally available model, exactly double the rate of Claude Opus 4.8.

The critical technical fact is this: token counts are computed server-side by Anthropic’s inference engine and are not independently observable by the customer. Unlike electricity consumption, which is metered by an independent device at the point of use, or data bandwidth, which can be measured by network monitoring tools, token consumption cannot be independently verified by a third party without provider cooperation. Vaudit’s SDK works by installing inside a customer’s environment and comparing provider-reported usage data against provider-reported billing data — not against an independent measurement. When the two figures match, the customer is being billed consistently. When they diverge, there is an overcharge. But in neither case can the customer confirm that the underlying count is accurate.

“What we are observing is that enterprise AI billing has become increasingly opaque,” Hahn said. “Customers often don’t have independent visibility into which model actually handled a request, how it was routed, or whether it was cached, retried, or deduplicated before it showed up on their bill.”

The billing chain grows longer when customers access Claude through cloud intermediaries — Amazon Web Services Bedrock, Google Vertex AI, or Microsoft Azure AI. A customer paying a cloud provider for Claude access may find the invoice has traveled through multiple hands before arriving, and that disputes require escalating through two companies rather than one.

Class Action Challenges Whether Max Subscriptions Deliver What Anthropic Advertises

The enterprise overcharge problem described by Vaudit is a billing system accuracy problem. A separate legal challenge filed June 14 in the U.S. District Court for the Northern District of California — docket 3:26-cv-05763 — argues a different claim: that Anthropic’s subscription marketing is false.

Washington, D.C. resident Karl Kahn, who filed the proposed class action, had progressively upgraded through Anthropic’s subscription tiers — Claude Pro in June 2025, Max 5x at $100 per month in January, and Max 20x at $200 per month this April — to support intensive coding work. According to the complaint, the Max 20x plan is marketed as delivering twenty times the usage capacity of the base Pro plan. Kahn alleges it delivered somewhere between six and eight times Pro usage in practice; the Max 5x plan delivered roughly three-and-a-half times, not five. One five-hour programming session consumed 15 percent of his entire weekly allotment. He found himself purchasing additional usage credits on top of a $200 monthly subscription.

The lawsuit alleges Anthropic “did not clearly explain how usage was measured, making it difficult for subscribers to determine whether they were receiving the amount of access promised.” The complaint also alleges the company “entices consumers to pay the $200 monthly Max 20x subscription by falsely claiming it offers a 50% savings.” Anthropic has declined to comment on the filing. No class has been certified, and no finding of wrongdoing has been made.

The measurement gap at the center of Kahn’s case is technically specific: Anthropic’s Max plans are governed by tokens and marketed by “session” multipliers — Max 20x, Max 5x — without publishing a precise per-session token budget. Because the company has never defined what constitutes a “session” or disclosed the token allowance underlying each multiplier, a subscriber has no mechanism to confirm whether a “20x” plan is actually delivering twenty times the throughput of Pro. That opacity is not incidental to the lawsuit — it is its structural foundation.

Kahn’s attorney, Kati Daffan of Vaca Daffan LLP, described the case in plain terms. “The law is very clear that companies have to be honest about their advertising and marketing. If they’re not, they can be held liable, need to change their ways, and refund money to people.”

A significant legal note: Anthropic’s Consumer Terms of Service — the terms governing personal Max subscriptions — contain no mandatory arbitration clause and no class action waiver. The Kahn lawsuit proceeds in federal court under California’s Consumers Legal Remedies Act and False Advertising Law.

A Year of Billing Incidents the Draft Never Had to Invent

The phantom invoice and the class action are the most visible points on a billing controversy that has been building since January. In late January 2026, a single user was charged 26 times in a single day; in February, double billing was confirmed as a system bug. Between March and April, at least six users across Taiwan, Poland, the United Kingdom, and the United States reported being charged for subscriptions they never purchased — including “Gift Max 5X” and “Gift Pro” subscriptions appearing on accounts without their authorization. Trustpilot reviews gave Anthropic a 3 out of 10 rating during this period, with billing complaints as the dominant theme.

The most technically revealing incident before July was the HERMES.md bug, disclosed in late April. On April 4, Anthropic began enforcing a policy that blocked third-party agent harnesses — software tools like OpenClaw and the open-source Hermes Agent from Nous Research — from drawing on Claude Code’s flat-rate subscription quota, requiring those workloads to be billed at pay-as-you-go API rates instead. The enforcement mechanism was a content classifier embedded in Claude Code that scanned git commit history for strings associated with those harnesses. The bug: the classifier was matching on the string “HERMES.md” appearing anywhere in recent commit messages — not on evidence of actual harness usage. A developer who had ever typed that filename in a commit message would find sessions silently rerouted to extra-usage billing, generating charges while their flat-rate subscription quota sat unused.

User sasha-id was charged $200.98 in extra usage credits while 86 percent of his Max 20x weekly plan remained untouched. When he contacted Anthropic support, the company initially refused a refund, describing the episode as an “un-refundable technical error.” Anthropic’s consumer support is primarily handled by an AI chatbot with no documented escalation path to a human agent; multiple users across this period reported the same pattern of the chatbot acknowledging billing errors in writing while no human ever followed up. The refund came only after the Reddit post about the bug reached approximately 1.44 million views and the Hacker News thread reached 1,251 points. Boris Cherny, Anthropic’s Head of Claude Code, acknowledged the issue as an “overactive anti-abuse system” and marked it fixed.

The pattern the byteiota.com analysis summarized applies here: support responses arrived at the speed of public pressure, not at the speed of the error.

Competitive Pressure and a Pending IPO Raise the Stakes

Anthropic’s billing problems arrive as the competitive landscape for AI pricing grows more aggressive. Microsoft unveiled its MAI-Thinking-1 model at its Build developer conference in June, and the company’s AI division chief Mustafa Suleiman told Bloomberg that “a lot of people are urgently looking for alternatives to Anthropic’s expensive models.” OpenAI is reportedly weighing significant token price reductions. DeepSeek V4 offers output token pricing roughly 90 times cheaper than Claude Sonnet 4.6 by some developer estimates, though DeepSeek operates under Chinese jurisdiction, where the National Intelligence Law (2017) and the Data Security Law (2021) legally compel the company to share user data with the government on demand — an obligation that applies regardless of where data is stored or what privacy policies the company publishes.

The Fable 5 arc added its own strain. Anthropic’s most capable publicly available model launched June 9, was suspended on June 12 by a US government export control directive requiring restrictions on foreign national access, and returned July 1 after the Commerce Department lifted those controls. The company offered a subscription-included window for paying users through July 7, then extended it to July 12 — tonight at 11:59 PM PT, the free window closes and Fable 5 moves to usage credit billing at $10 per million input tokens and $50 per million output tokens. That sequence — launch, government suspension, compressed free window, and usage-credit shift — added pricing uncertainty on top of the billing confidence problems already accumulating.

Anthropic is widely expected to pursue a public offering, and investor diligence will include user trust metrics alongside revenue. For enterprise finance teams now examining AI line items that can shift dramatically month to month, the question of whether a bill accurately reflects actual usage is a procurement and audit concern as much as a developer complaint. Zylo’s 2026 SaaS Management Index found that 78 percent of IT leaders reported unexpected charges from consumption-based AI pricing models. Ninety percent named AI cost forecasting as their top deployment challenge, according to Flexprice research. Uber reportedly burned through its entire 2026 AI budget for Claude Code in four months after rolling out access to roughly 5,000 engineers.

The emergence of firms like Vaudit — a 30-person startup that charges one percent of invoices reviewed plus 30 percent of amounts recovered, and has now audited more than $1.2 billion in total enterprise vendor spend — is itself a diagnostic. It indicates that AI billing has grown complex enough that major enterprises cannot independently verify what they owe. A market where the only check on billing accuracy is a third-party auditor charging a contingency fee is a market where trust in the underlying provider’s numbers has materially eroded.

What Customers Can Do Now

For anyone using Claude at any tier, the practical steps are concrete. Set hard monthly spending caps in the API dashboard before any new workload touches the system; without a cap, a misconfigured agent can run up charges before anyone notices. Enable usage alerts at meaningful thresholds — 50 percent, 75 percent, 90 percent of monthly budget — so a billing anomaly triggers an alert rather than an invoice. Keep independent logs of API call volumes and session counts; even if token counts cannot be independently verified, request counts can, and a discrepancy between your log and the provider’s count is a signal to dispute. If on a Max subscription plan, screenshot the dashboard weekly during heavy usage periods and retain those screenshots; in any dispute, time-stamped evidence of dashboard figures is the primary asset.

For enterprise buyers, the Vaudit findings suggest a practical threshold: any organization spending more than $100,000 per year on AI APIs should implement independent usage logging before the next billing cycle and reconcile it against provider invoices at least quarterly. Providers including Anthropic, Amazon, Google, Microsoft, and OpenAI refunded roughly 80 percent of properly challenged disputed amounts within 48 to 96 hours — the challenge is identifying what to dispute. That identification is what independent monitoring enables.

Whoever makes verification easy for customers — through transparent per-request model-routing logs, published session-token budgets, and an auditable billing API — will carry a structural advantage into the next phase of enterprise AI adoption. Billing transparency is becoming its own competitive battleground, and the companies that address it proactively will earn something more durable than a benchmark win: the confidence of a customer who does not feel the need to hire an auditor.

Frequently Asked QuestionsHow do I dispute an Anthropic billing error?

Document the discrepancy with screenshots of your usage dashboard showing the period in question, then submit a support ticket with the specific dates, dollar amounts, and what your own records show. Vaudit’s audit of 60 enterprise clients found that roughly 80 percent of formally disputed overcharges were credited back by Anthropic and other providers within 48 to 96 hours of being clearly presented. For free-tier users who receive erroneous invoices — as in the $16.6 million phantom invoice case — document the invoices and contact your bank to block unauthorized charge attempts while pursuing a support resolution. Anthropic’s consumer support is currently AI-mediated; expect slow escalation to a human agent and escalate via a credit card dispute if the initial response is a refusal.

Why can’t customers independently verify what they’re being charged for AI token usage?

Token counts are computed server-side by Anthropic’s inference engine and are structurally unobservable by the customer. Unlike electricity consumption or data bandwidth — both of which can be independently measured — AI tokens are an internal unit the model calculates during inference. Customers cannot intercept and count their own tokens without provider-supplied tooling or a proxy layer. Third-party billing auditors like Vaudit work by comparing provider-reported usage data against provider-reported billing data, which can reveal discrepancies but cannot confirm the accuracy of the underlying count. This structural asymmetry is the engineering root of the opacity problem — it is not a transparency choice that providers could simply reverse with a policy change.

What is the Kahn v. Anthropic class action about, and who qualifies?

Filed June 14, 2026, in the U.S. District Court for the Northern District of California (docket 3:26-cv-05763), the complaint alleges that Anthropic’s Claude Max 5x ($100/month) and Max 20x ($200/month) subscription plans deliver significantly less usage than their names advertise. Plaintiff Karl Kahn alleges Max 20x delivers roughly six to eight times the usage of a standard Pro plan, not twenty times, and Max 5x delivers roughly 3.5 times, not five. The proposed class covers all U.S. residents who purchased or upgraded to a Max 5x or Max 20x plan through Claude.com or the Claude desktop app since April 2024. No class has been certified and no finding of wrongdoing has been made. Notably, Anthropic’s Consumer Terms of Service do not contain a forced arbitration clause or class action waiver — the case proceeds in open court.

What is a “retry storm” in AI billing, and how does it generate unexpected charges?

A retry storm occurs when an autonomous AI agent — software designed to complete multi-step tasks by calling the Claude API repeatedly — encounters a failure in one step and automatically retries it. Each retry is a separate API call with a separate billing event. An agent caught in an infinite loop, or an agent set to retry aggressively without a maximum retry count, can generate hundreds or thousands of billable calls in minutes. Vaudit identified retry storms as one of the most common sources of enterprise AI overcharges in its audit of $34 million in invoices. The fix is architectural: set maximum retry limits and total cost caps in any agent loop before deployment, and enable hard API spending limits in the provider dashboard so that a runaway agent cannot exceed a defined monthly ceiling.