AI agents are dramatically boosting enterprise productivity, yet the companies adopting them are groaning under skyrocketing token usage fees. In response, AI service providers are competitively rolling out new service models to offset the cost burden, shaking up the landscape.

According to the information and communications technology (ICT) industry on July 5, AI companies are recently pursuing a multi-pronged approach to ease customer cost pressures — shifting to outcome-based billing, launching cost-effective new models, or unveiling technology that fundamentally reduces token consumption.

No Results, No Charge: Salesforce Upends the Billing Model

The most disruptive move comes from Salesforce (CRM). The company recently introduced “resolution-based pricing” for its Agentforce Help Agent, a service that supports building and operating customer service AI agents. Under this model, enterprises pay not based on the number of times an agent is called or usage volume, but only when the AI autonomously resolves a customer’s issue from start to finish. If a customer requests to speak with a human agent or expresses dissatisfaction with the outcome, no charge is incurred.

A Salesforce representative explained, “Enterprises will be able to improve AI investment efficiency based on actual outcomes rather than simple agent response counts.”

Anthropic Competes on Value; Microsoft Unveils Tech to Cut Tokens by 98%

AI startup Anthropic is countering with price competitiveness. On June 30, the company launched its new model, Claude Sonnet 5, stating it “delivers performance close to the higher-tier Claude Opus 4.8 at a lower price.” At promotional pricing, it costs $2 per million input tokens and $10 per million output tokens — just 40% of the Opus 4.8 pricing of $5 and $25, respectively.

Microsoft (MSFT) is focusing on fundamentally reducing token consumption through a technical approach. On June 29, Microsoft Research unveiled Memora, a memory architecture that improves long-term memory for AI agents. Currently, AI agents must reread lengthy conversation histories from scratch in every session, causing token usage to grow exponentially. Memora separates how information is stored from how it is retrieved, selectively recalling only necessary memories. Microsoft reported that in benchmark tests, Memora reduced token usage by up to 98% compared to existing methods while achieving higher accuracy. However, this remains a research achievement rather than a commercial product, so whether it will directly translate into lower bills remains to be seen.

‘FinOps’ for AI Spending Management Surges; OpsNow Enters the South Korean Market

Solutions for systematically managing token costs — known as “AI FinOps” — are also drawing attention. FinOps, a concept combining Finance and DevOps, originated as a methodology for optimizing costs by tracking cloud usage. It has recently evolved into AI spend management, encompassing GPU usage and large language model (LLM) call fees.

In the global market, companies like CloudZero, Vantage, and Finout offer services that aggregate usage costs from multiple AI model providers in a single view. In South Korea, OpsNow, an affiliate of Bespin Global, is expanding its footprint by supporting solutions through OpsNow FinOps that convert volatile AI spending into predictable operating costs.

“Token Costs at 30% of Payroll”: The Silicon Valley AI Economics Paradox

Driving these shifts is the very real cost burden facing enterprises. According to an internal analysis by SemiAnalysis, a leading semiconductor research firm in Silicon Valley, its internal LLM token spending has reached 30% of total employee payroll. That figure is more than five times the average per-employee usage at Meta. Yet SemiAnalysis assessed that “this cost handles tasks that previously required several times the headcount in just minutes,” describing it not as a mere efficiency gain but as “a phenomenon where the unit economics of professional services are being fundamentally restructured.”

In fact, Nvidia (NVDA) CEO Jensen Huang declared at this year’s GTC conference that “a $500,000-a-year engineer’s annual token cost should be no less than $250,000,” pledging to equip 75,000 employees with 7.5 million AI agents.

However, alarm bells over AI spending management are ringing across Silicon Valley. Uber (UBER) encouraged engineers to use Claude Code last year, only to see its annual budget exhausted within months. Usage had exploded, with 95% of engineers using AI monthly and 70% of all code submissions originating from AI. Uber ultimately responded by setting a monthly token usage cap of $1,500 per employee, requiring special approval for any overage.

The situation is similar at Microsoft. According to tech publication The Verge, Microsoft is canceling most Claude Code licenses and transitioning to its own GitHub Copilot CLI to control costs. Nvidia’s Vice President of Applied Deep Learning Research, Bryan Catanzaro, also lamented that his team’s compute costs had “far exceeded” employee costs.

The Paradox: Falling Costs, Rising Margins

The prevailing industry view is that AI token costs will decline over the long term. SemiAnalysis projects that software optimization alone on the latest Nvidia B300 and GB300 NVL72 systems can boost throughput by up to 14 times. For AI agents specifically, the input-to-output ratio reaches 300:1 with cache hit rates exceeding 90%, meaning the actual blended cost falls to roughly one-quarter of the list price.

This “cost decline” is paradoxically driving improved profitability for AI companies. Anthropic’s annual recurring revenue (ARR) has surged from $9 billion this year to over $44 billion, while its gross margin has jumped from 38% to above 70%. Even as token prices fall, exploding usage volumes are creating a structure where sellers reap greater profits.

Market research firm Gartner forecasts that by 2030, the inference cost for large-scale models with one trillion parameters will drop by more than 90% compared to 2025. Global tech companies’ AI capital expenditure this year has surged 69% year-over-year to $740 billion, yet the pace of tech industry layoffs has already surpassed last year’s levels. Before the full economic shock of AI materializes, managing the transitional “cost shock” has emerged as a critical challenge for enterprises.