As FinOps AI strategies continue to emerge,the familiar cloud cost management approach is breaking down — and organizations that fail to adapt risk runaway spending on workloads they barely understand.
The discipline of FinOps is rapidly evolving from a cloud-billing function into a strategic framework for governing the full technology stack, including AI, software as a service and now autonomous agents. According to the “State of FinOps 2026 Report,” 98% of practitioners now manage AI spend — yet most organizations still lack the cost granularity needed to govern it effectively, according to Pravir Gupta (pictured), vice president and general manager of Google Cloud at Google LLC.
“The same trend will continue,” Gupta said. “Every CEO is asking their teams, the whole organization to innovate fast with gen AI, and that’s where FinOps is still very relevant — to make sure that you have the right guardrails to better estimate the cost, to have explainability of those costs, as well as the guardrails that allow you to innovate faster.”
Gupta spoke with theCUBE’s John Furrier and Paul Nashawaty at FinOps X 2026, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed the future of FinOps AI, tokenomics, agentic cost structures and how Google applies generative AI internally to drive measurable business transformation. (* Disclosure below.)
FinOps AI demands cost granularity beyond tokenomics
Token economics is a centerpiece of the FinOps AI cost conversation, but treating it as the complete answer misses most of the picture. When an AI agent executes a task, it may also spin up virtual machines, consume key-value cache storage and trigger retrieval-augmented generation pipelines — costs that sit entirely outside the input-output token line item, Gupta noted.
“Tokenomics is a large piece of the FinOps for AI, but the real thing for every enterprise to focus on within the FinOps area is the FinOps for AI,” Gupta said. “It’s like the iceberg — what’s under the water. There are input tokens and output tokens, but the agent may spin up a VM in a sandbox to actually write scripts and do things. You may also have what you call adjacent AI costs from your key-value cache … All of those are additional adjacent costs outside just the input and output token.”
Google itself has demonstrated what rigorous AI cost accountability can unlock in practice. As customer zero for its own platform, Google’s business transformation program — internally called Google on Google AI — applied an orchestrating agent to supplier invoice reconciliation across all of Alphabet Inc., shifting humans from doing to reviewing the output of agents. The result was a four times increase in throughput capacity and $30 million in savings, Gupta said.
“That same pattern applies in so many different ways because the trick here was not to roll out with a hundred percent accuracy,” he said. “The trick is that you have a human in the loop in the middle where humans are reviewing the output of the agent and then providing the feedback.”
As headless agents become more prevalent, cost attribution grows even more complex. Gemini Spark, Google’s newly announced 24/7 personal agent for Workspace, exemplifies where the industry is heading: orchestrator agents that initiate workflows autonomously and call sub-agents, each potentially running on a different model tier. Governing that cost structure requires granularity at every layer — by orchestrator, by sub-agent, by model, and by organizational tag — so that chargeback and anomaly detection remain meaningful as agentic work scales.
“You need the cost granularity at all the different dimensions,” Gupta said. “I need to know not just the cost of the orchestrator agent, but the sub-cost of each of the underlying agents, as well as the cost by input, output token, the cost by different models — because each agent may not need to use the same model. If it’s running a smaller task, it can use a flash, a workhorse model, whereas a more complex task can use your frontier models. So the cost granularity becomes really important so that you have cost explainability of that overall.”
Here’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of FinOps X 2026:
(* Disclosure: TheCUBE is a paid media partner for the FinOps X event. Neither the FinOps Foundation, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)
Photo: SiliconANGLE
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.
About SiliconANGLE Media
SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.