{"id":119243,"date":"2026-07-26T11:54:11","date_gmt":"2026-07-26T11:54:11","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/119243\/"},"modified":"2026-07-26T11:54:11","modified_gmt":"2026-07-26T11:54:11","slug":"modelmaxxing-replaces-tokenmaxxing-for-firms-grappling-with-ai-costs","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/119243\/","title":{"rendered":"\u2018Modelmaxxing\u2019 Replaces \u2018Tokenmaxxing\u2019 for Firms Grappling With AI Costs"},"content":{"rendered":"<p class=\"wp-block-paragraph\">To borrow from the British idiom, companies no longer want to use an AI sledgehammer to crack a nut.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">It seems like just yesterday that ChatGPT came on the scene, changing everything from how people proofread their resumes and make grocery lists to the way they plan vacations. As more tools like Claude and Gemini entered the picture, so did so-called \u201c<a href=\"https:\/\/www.nytimes.com\/2026\/03\/20\/technology\/tokenmaxxing-ai-agents.html\" rel=\"nofollow noopener\" target=\"_blank\">tokenmaxxing<\/a>\u201d in which employees competed, often at the behest of managers, to use the most tokens on the job. (Tokens are the units used to measure AI use and determine payment.)\u00a0<\/p>\n<p class=\"wp-block-paragraph\">But pulling out Anthropic and OpenAI\u2019s offerings for basic tasks can be akin to calling an Uber just to go around the block or an electrician to change your light bulb. The task can be done much more simply, and for a lot less money, another way. Enter \u201c<a href=\"https:\/\/www.businessinsider.com\/ai-model-routing-modelmaxxing-efficient-token-use-2026-7\" rel=\"nofollow noopener\" target=\"_blank\">modelmaxxing<\/a>,\u201d which entails routing tasks so that they\u2019re done by the most efficient and cost-effective model.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe AI race is no longer just about building the best model,\u201d said Soumen Mandal, principal analyst at Counterpoint Research. \u201cWhile performance will remain important for both enterprises and consumers, the right balance of capability, cost and deployment flexibility will ultimately determine the winners in the AI market.\u201d\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Those winners increasingly include Chinese AI companies that are offering capable models for a fraction of the cost of their US rivals. Just this month, Chinese AI firm Moonshot <a href=\"https:\/\/www.thedailyupside.com\/technology\/artificial-intelligence\/moonshot-alibaba-place-ai-affordability-in-the-spotlight\/\" rel=\"nofollow noopener\" target=\"_blank\">launched Kimi K3<\/a>, a model that outperformed all rivals except Anthropic\u2019s Claude Fable 5 and OpenAI\u2019s GPT-5.6 just before Alibaba previewed its Qwen3.8 Max model, which it claimed is even more powerful.<\/p>\n<p>Model Money Management<\/p>\n<p class=\"wp-block-paragraph\">In June, Coinbase CEO Brian Armstrong <a href=\"https:\/\/x.com\/brian_armstrong\/status\/2070670644577280109?lang=en\" rel=\"nofollow\">took to X<\/a> to explain how to keep AI spend flat while token usage grows, including by better routing. \u201cIn our custom harnesses, we preprocess prompts and route to the best model for the job, considering cache hits and model pricing,\u201d Armstrong wrote.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">It\u2019s no surprise that cost is top of mind for companies now that there are options for their AI. \u201cThis is the first year where AI spend has been a big topic,\u201d OpenAI CEO Sam Altman told <a href=\"https:\/\/www.cnbc.com\/2026\/07\/09\/cnbc-exclusive-transcript-openai-ceo-sam-altman-speaks-with-cnbcs-julia-boorstin-on-squawk-on-the-street-today.html\" rel=\"nofollow noopener\" target=\"_blank\">CNBC<\/a> at the annual Allen &amp; Co. Sun Valley Conference. \u201cAnd all of a sudden, it\u2019s a very big topic. Everyone\u2019s asking what we can do to help reduce spend or increase value.\u201d\u00a0<\/p>\n<p class=\"wp-block-paragraph\">As companies consider their options, Chinese large language models such as DeepSeek, Doubao, GLM, Hunyuan, Kimi and Qwen have made major strides in the US. An IDC survey shows that 47% of 260 US decision-makers at companies with more than 1,000 employees reported using made-in-China models for at least one use case, and 20% reported extensive use. In a separate survey from late last year of roughly 100 respondents, 73% reported using or testing AI-driven routing, and 72% reported using or testing automated routing based on input or metadata analysis.<\/p>\n<p class=\"wp-block-paragraph\">While the firm can\u2019t conclude from the data that the demand for model selection and routing is primarily driven by cost savings, the pricing advantage of Chinese AI models is substantial. Chinese models like DeepSeek and Qwen typically charge less than $5 per million output tokens while US frontier models like GPT 5.5 and Claude Opus are priced about $25 to $30, Mandal said.\u00a0<\/p>\n<p>System Rerouting <\/p>\n<p class=\"wp-block-paragraph\">Model routing is becoming a more important part of the AI stack, said Adam Swick, head of strategy at OpenRouter, a marketplace for AI models. A common pattern is to use a highly capable model to plan a complex task, break it into steps and determine what needs to happen, then use lower-cost models to execute the more routine parts of that plan.<\/p>\n<p class=\"wp-block-paragraph\">\u201cWe have seen meaningful growth in open-weight models, particularly for coding, agentic workflows, reasoning and high-volume applications where cost efficiency matters,\u201d Swick said. (Open-weight models, like those released by DeepSeek, keep their core components public, while closed-weight ones like those from Anthropic and OpenAI don\u2019t.) \u201cDevelopers are increasingly evaluating models based on how well they perform for a specific workload rather than relying solely on the lab or brand behind them.\u201d\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Closed model options tend to be more accurate, but more expensive and slower. Open-weight ones can be less accurate, but they\u2019re also fast, cheap and offer developers more control. Plus, with a specialized open-weight model, an enterprise can get a product that is fast, cheap and accurate for a specific task, said Iz Beltagy, research scientist at the Allen Institute for AI (Ai2).\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cIt\u2019s true that this model won\u2019t be good at everything else, but who cares if it is good at the one task they care about?\u201d Beltagy said.\u00a0<\/p>\n<p>Risky Business<\/p>\n<p class=\"wp-block-paragraph\">But companies have more to think about than cost. There\u2019s also a security concern. The US recently <a href=\"https:\/\/www.reuters.com\/world\/asia-pacific\/pentagon-lists-entities-designated-chinese-military-company-2026-06-08\/\" rel=\"nofollow noopener\" target=\"_blank\">added<\/a> Chinese e-commerce company Alibaba, internet search provider Baidu and other tech giants to a list of companies it says are aiding the Chinese military. The US Defense Department won\u2019t be allowed to contract with the firms.<\/p>\n<p class=\"wp-block-paragraph\">\u201cBeing public about using a Chinese LLM in the US market triggers, at best, suspicion; at worst, reputational damage and procurement pushback,\u201d Igor Marchal, a vice president analyst at Gartner wrote in a recent report. \u201cOperationally, if you do not lock down how a model runs, you risk leaking sensitive corporate data or siphoning intellectual property into untrusted infrastructure.\u201d\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Swick said it\u2019s important to distinguish between the model lab that develops a model and the provider that actually hosts and serves it from a data center. Data handling, security, compliance, geography and reliability can depend heavily on the provider and deployment architecture, not simply on where the underlying model was originally developed, he added.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Many developers pull \u201cfine-tuned derivatives\u201d from open-source repositories, Marchal wrote, not knowing that the underlying architecture was built on a Chinese model base and \u201cintroducing compliance and shadow AI potential headaches.\u201d\u00a0<\/p>\n<p>The Performance Gap\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The US is still ahead when it comes to figuring out where technology is heading, pushing the frontier and answering the really difficult questions, Beltagy said. (All the major capabilities, such as coding and agents, were introduced by US frontier labs and reproduced by everyone else.) The US also has a unique strength in fully open AI, not only in releasing models, but also in the training data, code, checkpoints, evaluations and methodology, so that anyone can understand, reproduce, and build on the work, he added. That allows the continued building of a vibrant ecosystem that\u2019s not locked behind closed doors.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">But Chinese models have made significant improvements in recent years in coding, multilingual capabilities, mathematics and productivity tasks, helping narrow the performance gap with leading US frontier models, Mandal said.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cImprovements in model efficiency without sacrificing quality, the availability of open-weight models and government support for sovereign AI development are helping Chinese AI companies achieve economies of scale,\u201d he added.\u00a0<\/p>\n<p>What\u2019s Next?\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The fact that tokenmaxxing made headlines just months ago and is already being replaced by the next trend shows just how fast things move in the world of AI models.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In the future, Mandal said AI models will be commoditized for standard workloads, and long-term differentiation will come from how AI is deployed rather than from the models themselves. The value proposition will shift toward applications, enterprise software and infrastructure.<\/p>\n<p class=\"wp-block-paragraph\">That means the next phase of competition will be defined by commercial viability (as in the balance of performance, cost, capability and deployment flexibility) rather than benchmark performance.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cSovereign AI investments will reshape the competitive landscape, while the multi-model approach will create regional leaders rather than a single global winner,\u201d he added. \u201cAs AI adoption is still in its early stages, we expect value creation to extend beyond foundation models to cloud providers, semiconductor companies, enterprise software vendors and AI application developers.\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"To borrow from the British idiom, companies no longer want to use an AI sledgehammer to crack a&hellip;\n","protected":false},"author":2,"featured_media":119244,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,2477,321,53,25,387,132,1122,30971,157],"class_list":["post-119243","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-alibaba","tag-amazon","tag-anthropic","tag-artificial-intelligence","tag-china","tag-google","tag-meta","tag-moonshot","tag-openai"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/119243","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=119243"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/119243\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/119244"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=119243"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=119243"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=119243"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}