
How LLM-agnostic applications are liberating the enterprise from single-vendor lock-in with frontier labs.
Claude, Gemini, F2
When OpenAI’s Codex and Anthropic’s Claude Code arrived on the scene, companies found a new way to compete. Not by revenue growth or market share, but token use, known as “tokenmaxxing.”
The idea was that tokenmaxxing meant employees were multiplying their output with AI. And as eager as employees were to outsource their work to agents, managers and executives were right there with them.
Token use was incentivized from the top down, often at the urging of FOMO-stricken boards. Meta and Amazon ran employee token-use dashboards to drive competition among engineers, while Coinbase went as far as to fire technical staff who didn’t adopt AI fast enough.
At first, the costs were manageable. Companies were adding a new line item, but in the name of discovery, finding where this seemingly infinite new resource could create real value. Finance teams were cautiously watching the invoices roll in.
Then the economics changed.
Each new flagship model reasoned longer and burned more tokens on the same tasks, just as employees got creative, steering the coding assistants with their own autonomous agents. The labs responded by shifting autonomous agents and third-party tools from flat-rate subscriptions to metered, pay-per-token credits, and the line item earmarked for discovery became one of the fastest-growing costs in the corporate budget.
Take Uber, for example, which burned through its entire 2026 budget for AI coding tools by April, and is now capping each employee at $1,500 per tool per month in token spend; or ServiceNow, which blew through its full-year Anthropic budget in the first few months of 2026. Even a Mag 7 company like Microsoft canceled most of its internal Claude Code licenses in May and moved its engineers to GitHub Copilot.
The rise in token costs among frontier models was originally the entire argument for multi-model systems. Now, we see it was just the first domino to fall. When companies build their core workflows on a single frontier lab’s models, they become subject to the lab’s ability to decide when, how, and for what reason they’re using their models. The more you embed your systems into a single lab’s models, the harder it becomes to unravel your workflows and the more expensive it becomes to switch.
Frontier Labs controlled the price and access. Until now.
On June 15th, Anthropic is moving programmatic Claude usage onto a separate metered credit, billed at full API rates. Workflows that ran under a flat monthly subscription now draw down credits, by some estimates, a 12x to 175x effective price increase, depending on the task.
And models are only getting more token-dense. Claude Fable 5, the latest top-end model, launched on June 9 at $10 per million input tokens and $50 per million output tokens — double the price of Opus 4.8 and the most expensive major model on the market. It was originally intended to be included in subscriptions before moving to the same metered pricing on June 23rd. That was until the model was pulled from production (more on this later).
We can see this leverage reflected in Anthropic’s historic growth. The company closed a $65 billion round at a $965 billion valuation in late May with run-rate revenue at $47 billion, up from $9 billion at the end of last year. With a probable IPO later this year, the company has every incentive to protect the revenue on which its valuation depends, and none to cut prices on the way up. For anyone wondering whether high token costs are permanent or a blip, I think we have our answer.
While some of that growth was earned net revenue retention and new business, much of it likely comes from enterprises that were built around a single model, assumed token costs would stay low, and are now facing an exploding operating expense they never modeled.
Beyond price changes, companies are also at the whim of new terms and control provisions. Just in the past week, Anthropic launched its most capable public model, Fable 5, and then a government export-control order forced the company to immediately pull Fable 5 from every customer. The same model required 30 days of data retention, with no zero-data retention option even for enterprises that hold one on every other model, and added a safeguard that reroutes questions on cybersecurity, biology, and frontier AI work to a less-capable model. Each change was implemented with little warning, and any internal workflow built on the latest model was affected by these factors.
Today, though, companies don’t have to accept that dependence. A lab’s control over price and access rests on scarcity, and scarcity is ending. The cost of running AI at a fixed level of capability has been falling by a factor of 9x to 900x per year, and Anthropic’s ~50% gross margins only exist when tokens are scarce.
The “capability gap” is closing. The “pricing gap” is not. Relative cost to run the same workload.
Claude, Gemini, F2
Fortunately for companies, token scarcity is nearly over because viable substitutes can now handle most enterprise workflows. Over the past year, the lag between the closed frontier and the best open-weight alternative has compressed from roughly 12 months to roughly 3 months, while open-weight tokens have cost 8 to 100 times less. When several providers can do the same work, no company has to tie its operations to any one provider.
Companies are realizing the vast majority of their tasks don’t require frontier intelligence. Most enterprise work is repetitive and structured. It needs reliability, permissions, and auditability. It rarely needs the most expensive model in the market.
Justine Moore compares using a blow torch to light a cigar to using Opus 4.8 to complete a commodity task.
X.com
Brian Armstrong wrote last week that demand for intelligence is “near infinite,” but he expects 80% of workloads to run on models that are 99% cheaper within 12 to 18 months, reserving the newest models for the hardest 20%. Coinbase already routes prompts to cheaper models where the task allows.
This will be the standard in enterprise software. Token usage will grow exponentially, costs will remain flat or decline, and the company will retain the final say over which model runs each task.
Where enterprises can find the off-ramp
When an enterprise is built around a single frontier lab, it faces concentration risk: a single pricing schedule, rate-limit policy, model availability, and privacy and security terms. The more of its core work runs through that one lab, the less of its own operation it actually controls.
The labs understand this, which is why they’re pushing customers to build their companies’ internal tools directly in Claude, Codex, or Gemini. But under the surface, this creates lock-in. Tools built natively inside one lab’s ecosystem concentrate token spend, are hard to rip out, and hand the lab lasting power over price, terms, and access.
Now that single-lab risk is top of mind, more attention is being given to the application layer, which enables companies to switch between model providers.
An enterprise application layer routes each task based on cost, risk, and complexity, and reserves the frontier model for the most intelligence-intensive tasks, while cheaper closed models or open-weight models handle the long tail of routine work.
When you pair model-agnostic routing with persistent, structured memory, routine work may stop burning tokens altogether. Pre-processing multimodal data enables these tasks to be far more efficient than a generic LLM chatbot. The system gets cheaper and faster as it stores more data, and the knowledge it accumulates stays in the company’s hands.
Enterprise procurement teams are already catching on. Large companies spent $37 billion on generative AI in 2025, with more than half ($19 billion) going to the application layer.
In that architecture, the model becomes an input and can be priced, benchmarked, and swapped like any other line item. Every new model release from any lab improves the system. And no single lab’s decision to reprice, restrict, or withdraw a model can break what you’ve built. Every workflow documented makes it cheaper. You keep the upside, and you keep control.
The independence era
Frontier labs are pushing the limit of what artificial intelligence can do. And it’s true, the hardest problems in science, engineering, and finance will demand the best models ever built.
But most enterprise work lives elsewhere. It’s routine, predictable, and well-defined. It deserves an architecture built for it: routing that matches each task to the cheapest model that does it well, memory that stops paying for the same answer twice, and independence from any single vendor that can pull the model out from under it.
Companies used to measure AI by how much of it they used. Leaderboards ranked teams by tokens burned, and a big bill meant you were serious. What matters now is whether your workflows can continue to run without depending on any individual model or frontier lab.
The companies that get there first will use more AI than ever, for less, and on their own terms. Freedom from the frontier labs. Freedom for the enterprise. FREEDOM!