Token prices are falling, yet enterprise AI bills keep climbing. ManageEngine’s Rajesh Ganesan says the fix is spending smarter, and the data is starting to agree.

Image:
Rajesh Ganesan, ManageEngine CEO
Rajesh Ganesan has a blunt read on where agentic AI costs are heading: most enterprises, he argues, are paying for power they will rarely use, on a pricing model that was never built to last. The ManageEngine CEO, in a recent media roundtable ahead of its Southeast Asia User Conference in Jakarta, calls token-based AI consumption “already not sustainable,” and says his company would rather build models efficient enough that it never has to charge for them by the token.
“It need not be very costly,” he says of enterprise AI, casting efficiency as a conviction the company held long before the current bills started landing. Coming from the enterprise IT arm of Zoho, that could be read as a vendor talking its own book. The problem for anyone inclined to dismiss it is that the numbers now agree with him.
The counterintuitive part is that AI is getting cheaper by the unit even as the bills climb. Ramp’s enterprise spending data shows the average cost of a million tokens fell from around $10 to $2.50 in a single year, and the per-token cost of intelligence has dropped by roughly 98% since early 2024.
Yet enterprise invoices keep rising, because the number that drives the bill is not price, it is volume. Gartner’s March analysis found that agentic workflows, which chain multiple model calls together to finish a task, burn between five and thirty times more tokens than a single chatbot query. Consultancy EY estimates that a customer-service interaction costing four cents in 2023 runs closer to $1.20 today once tools and reasoning loops are added, about thirty times more.
Goldman Sachs expects total token consumption to multiply roughly 24 times by 2030. The result is a run of budget blowouts at companies that are anything but careless. Uber’s chief technology officer said the budget he expected to need was gone before he could use it, after adoption of agentic coding tools jumped from about a third of its engineers to more than four-fifths between December and March, exhausting the annual AI budget by April.
Microsoft reportedly pulled most of its internal licences for one agentic coding assistant over runaway token bills. And in June, OpenAI’s Sam Altman conceded that the question of whether AI spending will pay off is “the most fair criticism right now of AI,” telling CNBC that customers had gone from never mentioning cost to raising it constantly.
Ganesan reads that pattern as the predictable cost of learning the hard way. “Most people first do things, get hit, fall down,” he says. “Only then they realise.”
The case for handling agentic AI costs more carefully
This is where Ganesan’s argument stops being a slogan. An analysis of 2.4 billion enterprise API calls found that routing every workload to a frontier model cost $18.40 per million tokens, while a tiered setup that sent routine jobs to cheaper models landed at $2.31, an eight-fold gap on work that mostly involves classification and summarisation rather than hard reasoning.
Most enterprise tasks, in other words, do not need frontier capability, and paying frontier prices for them is the largest controllable line in the bill. ManageEngine’s answer is a gateway it calls Platform AI, which lets a customer decide which model runs underneath a given workload. “It is flexible by default,” Ganesan told reporters, defaulting to Zia, the company’s own model, with heavier options such as Claude available when a task warrants them.
That is model routing by another name, and the same instinct runs through the bigger problem he sees coming. As enterprises adopt agents from multiple vendors, Ganesan puts token cost and consumption at the top of the list of things they will struggle to control, alongside the governance headache of assigning identities to each agent so they, in his words, “operate in a governed manner.”
ManageEngine runs agents inside its own walls already, he says, and still has to handle much of that oversight manually. That shift is where the channel comes in. Managing token spend has become a discipline in its own right, and one most enterprises are not equipped for. The share of FinOps practitioners responsible for AI spend jumped from 31% in 2025 to 98% in 2026, a function that spent a decade rightsizing cloud infrastructure now handed a cost structure with no established playbook.
The pricing ground is moving under them too: GitHub shifted its Copilot coding assistant to usage-based billing in June, ending an arrangement where heavy users quietly consumed several times the compute their subscription covered. Partners and managed service providers are the obvious answer to that gap, building the model-routing and cost-governance layers that clients have neither the tooling nor the visibility to run themselves.
The vendors that win will be the ones whose partners can actually operate that discipline for customers. None of this means the cost problem is permanent, and the honest version of the story admits as much. Paul Roetzer of the Marketing AI Institute argues that the price of any given model will fall ten to a hundred times within a year, which could make today’s crisis largely evaporate, and that appetite for intelligence is so deep the industry is still at “the top of the first inning.”
There is also an awkwardness in casting efficiency as an insight the big labs missed. Google used its 2026 developer conference to pitch a cheaper, faster model it said could save enterprises more than a billion dollars a year, urging customers to move most of their workloads off frontier models. Ganesan is betting that instinct spreads.
He expects “more companies, more people, more nations will agree with this philosophy of keeping AI efficient.” He may be right, which would make ManageEngine one of many arriving at the idea rather than the only one holding it.
Cheaper models might eventually make careful routing less urgent, but nobody staring at this month’s AI bill can count on that yet. The enterprises drowning in token spend need discipline now, and Ganesan knows one vendor cannot supply all of it. “We can’t do everything ourselves,” he says, betting the approach spreads because the alternative does not hold. “We keep asking, is it sustainable? It is not.”