{"id":136668,"date":"2026-08-12T01:43:16","date_gmt":"2026-08-12T01:43:16","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/136668\/"},"modified":"2026-08-12T01:43:16","modified_gmt":"2026-08-12T01:43:16","slug":"why-cpus-still-matter-in-the-age-of-ai-agents-2","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/136668\/","title":{"rendered":"Why CPUs still matter in the age of AI agents"},"content":{"rendered":"<p>When the conversation turns to AI infrastructure, it almost always lands on GPUs and TPUs. The New Stack sat down with <a href=\"http:\/\/linkedin.com\/in\/bhumikpatel\" class=\"ext-link\" rel=\"external  nofollow noopener\" onclick=\"this.target=&#039;_blank&#039;;\" target=\"_blank\">Bhumik Patel<\/a> of Arm and <a href=\"https:\/\/www.linkedin.com\/in\/mofarhat0\/\" class=\"ext-link\" rel=\"external  nofollow noopener\" onclick=\"this.target=&#039;_blank&#039;;\" target=\"_blank\">Mo Farhat<\/a> of Google to talk about the chip that rarely makes the headlines anymore: the CPU, and why it\u2019s getting more important, not less, as AI shifts from chatbots to agents.<\/p>\n<p><a href=\"https:\/\/www.linkedin.com\/in\/mfarhat\/\" class=\"ext-link\" rel=\"external  nofollow noopener\" onclick=\"this.target=&#039;_blank&#039;;\" target=\"_blank\">Farhat,<\/a> who leads product management for Axion and Arm-based virtual machines at Google Compute Engine, tells The New Stack, \u201cThe role, more or less, is of a CPU as an air traffic controller.\u201d<\/p>\n<p>In this episode, we discuss how the shift from conversational chatbots to autonomous agents is quietly turning into a CPU story.<\/p>\n<p>\u201cToday\u2019s six- to eight-billion-parameter models are performing much better than they have in the past,\u201d Farhat says. For some specialized workloads, he says, CPUs can deliver roughly 25 tokens per second, which can be enough for agentic workloads.<\/p>\n<p>The workload shifted from answering to acting<\/p>\n<p>Early chatbots returned a response, but agents can act on them. They perform tasks by calling tools and, when needed, create environments to execute the code they write.<\/p>\n<p>\u201cThe orchestration harnesses themselves for agentic workloads are these always-on branching kind of control-flow logic that CPUs are great at,\u201d Farhat says.<\/p>\n<p>While large language models typically run on accelerators, CPUs also handle orchestration, data preparation, semantic search, and vector databases, Farhat says.<\/p>\n<p><a href=\"https:\/\/www.linkedin.com\/in\/bhumikpatel\" class=\"ext-link\" rel=\"external  nofollow noopener\" onclick=\"this.target=&#039;_blank&#039;;\" target=\"_blank\">Patel<\/a>, who drives Arm\u2019s software ecosystem efforts for cloud and AI, says the company is focused on the software and infrastructure layers needed to run these workloads at scale. Different types of agents, he notes, are doing \u201cdifferent type[s] of code execution and API calling and the typical CPU work.\u201d<\/p>\n<p>There\u2019s a role for actually running models here, too, but we\u2019re talking about very small ones, including summarizers, recommenders, and evaluators. <\/p>\n<p>\u201cToday\u2019s six- to eight-billion-parameter models are performing much better than they have in the past,\u201d Farhat says. For some specialized workloads, he says, CPUs can deliver roughly 25 tokens per second, which can be enough for agentic workloads.<\/p>\n<p>Why agents need sandboxes, and lots of them<\/p>\n<p>For those agents to run code, though, they need an environment that lets them do so securely without endangering production systems.<\/p>\n<p>\u201cThe agents are doing code execution, so you want to make sure that the LLM-generated code is trusted,\u201d Patel says. \u201cOr if it\u2019s not trusted, then you kind of sandbox the environment.\u201d<\/p>\n<p>Google\u2019s pitch for this is gVisor, an open-source project that acts as an isolation layer between the application and the host operating system. As Farhat puts it, \u201cWe operate in a zero-trust environment. This is exactly why you need the isolation technologies that agents run in.\u201d<\/p>\n<p>As Farhat puts it: \u201cWe operate in a zero-trust environment. This is exactly why you need the isolation technologies that agents run in.\u201d<\/p>\n<p>GKE Agent Sandbox, Google argues, can also handle the scale necessary in this agentic era. <\/p>\n<p>\u201cGKE Agent Sandbox will allow customers to spin up 300 sandboxes per second per cluster,\u201d Farhat says. Patel says the platform uses pod snapshots and warm pools of suspended environments to help customers scale more quickly without keeping all of those environments fully provisioned.<\/p>\n<p>The efficiency pitch<\/p>\n<p>Google says Axion can offer advantages in both cost and energy use. <\/p>\n<p>\u201cAxion today will give you up to 2x the price performance of comparable current-generation virtual machines,\u201d Farhat says, \u201cand we\u2019ll do that at over 60% better energy efficiency as well.\u201d<\/p>\n<p>Google offers Axion C4A machines for consistently high performance and N4A machines for a balance of performance and cost-optimized workloads. For a compute-bound, long-running engineering job where completion time is critical, Farhat says C4A offers the needed performance. For smaller code-execution tasks where density and cost matter more, he points to N4A.<\/p>\n<p>\u201cThe good thing about the options in Google Cloud is, like Mo offered earlier, you have the C4A for high-performance workloads, and then the N4A,\u201d Patel says. \u201cYou want to scale a large number of agents, and different types of agents are doing different types of code executions and API calling and the typical CPU work. The N4A is a great fit.\u201d<\/p>\n<p>Farhat\u2019s broader point is also that CPU, GPU, and TPU resources will continue to work together as agents scale. <\/p>\n<p>\u201cWe\u2019re in a fluid compute world,\u201d he says. \u201cWe do recommend that customers look at all the options, CPU-based, TPU, GPU-based, and ensure that their application is built to scale going forward.\u201d<\/p>\n<p>\t<a class=\"row youtube-subscribe-block\" href=\"https:\/\/youtube.com\/thenewstack?sub_confirmation=1\" target=\"_blank\" rel=\"nofollow noopener\"><\/p>\n<p>\n\t\t\t\tYOUTUBE.COM\/THENEWSTACK\n\t\t\t<\/p>\n<p>\n\t\t\t\tTech moves fast, don&#8217;t miss an episode. Subscribe to our YouTube<br \/>\n\t\t\t\tchannel to stream all our podcasts, interviews, demos, and more.\n\t\t\t<\/p>\n<p>\t\t\t\tSUBSCRIBE<\/p>\n<p>\t<\/a><\/p>\n<p>    Group<br \/>\n    Created with Sketch.<\/p>\n<p>\t\t<a href=\"https:\/\/thenewstack.io\/author\/frederic-lardinois\/\" class=\"author-more-link\" rel=\"nofollow noopener\" target=\"_blank\"><\/p>\n<p>\t\t\t\t\t<img decoding=\"async\" class=\"post-author-avatar\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/04\/15a7eb12-cropped-4e88ac40-frederic-profile-2-600x600.jpg\"\/><\/p>\n<p>\n\t\t\t\t\t\t\tBefore joining The New Stack as its senior editor for AI, Frederic was the enterprise editor at TechCrunch, where he covered everything from the rise of the cloud and the earliest days of Kubernetes to the advent of quantum computing&#8230;.\t\t\t\t\t\t<\/p>\n<p>\t\t\t\t\t\tRead more from Frederic Lardinois\t\t\t\t\t\t<\/p>\n<p>\t\t<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"When the conversation turns to AI infrastructure, it almost always lands on GPUs and TPUs. The New Stack&hellip;\n","protected":false},"author":2,"featured_media":136669,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[405,7537,132,3229,2454],"class_list":["post-136668","post","type-post","status-publish","format-standard","has-post-thumbnail","category-agentic-ai","tag-ai-agents","tag-artificial-intelligence-agents","tag-google","tag-podcast","tag-video"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/136668","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=136668"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/136668\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/136669"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=136668"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=136668"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=136668"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}