Central to the expanded partnership is Google’s Arm‑based Axion processor, which is now positioned as the primary CPU platform within the company’s rack‑scale AI systems.

Google Cloud’s move places it alongside other major cloud providers that have adopted Arm architectures as a foundation for large‑scale AI workloads. The company says the decision reflects a broader industry shift towards purpose‑built CPU designs as agentic AI transitions from experimental projects to production‑grade systems.

Alongside the expanded use of Axion, Google Cloud has announced updates to its AI infrastructure portfolio, including new eighth‑generation Tensor Processing Unit (TPU) systems and the introduction of an Agent Sandbox capability within Google Kubernetes Engine (GKE). These developments are intended to address the growing infrastructure demands of agentic AI, which requires continuous orchestration of reasoning, tool execution and real‑time data access rather than isolated model inference.

Agentic systems place new pressures on compute infrastructure. Unlike traditional inference workloads, they involve long‑running, multi‑step processes with high levels of concurrency and sensitivity to latency. As a result, CPUs play a more critical role in overall system performance, handling coordination, data preparation and execution control alongside specialised accelerators.

Google’s Axion processor, based on Arm Neoverse technology, has been designed to meet these requirements by delivering high throughput and energy efficiency at scale. Its integration as the host CPU in Google’s latest TPU platforms marks a departure from previous generations. The new TPU 8 family introduces separate variants for training and inference, with Axion serving as the central processing layer to reduce data preparation latency and keep accelerator resources fully utilised.

The infrastructure strategy also forms part of Google Cloud’s broader “AI Hypercomputer” vision, which combines custom silicon, open software frameworks and globally distributed infrastructure. A key element of this approach is the GKE Agent Sandbox, a deployment framework designed to allow AI agents to execute untrusted code and external tool calls securely and at low latency.

Running on Axion processors and built on gVisor with support for Kata Containers, the Agent Sandbox is designed to scale rapidly while maintaining fast startup times. Google Cloud says the environment can support hundreds of sandboxes per second within a single cluster while maintaining sub‑second time‑to‑first‑instruction latency. This level of performance, the company notes, is essential as agent‑based architectures become more widely adopted in production environments.

Efficiency at scale is also a central consideration. As agentic systems expand, the cost and power consumption of inference become increasingly important. By consolidating orchestration and inference workloads on CPU‑based platforms, Google Cloud aims to balance performance with operational efficiency, ensuring that accelerators are used effectively without excessive overhead.

Within the Axion family, Google Cloud offers several instance types tailored to different workload profiles. C4A virtual machines, powered by Arm Neoverse V2‑based Axion CPUs, are designed to complement accelerator‑centric systems by handling parallel and latency‑sensitive tasks such as AI inference and data processing on general‑purpose compute.

For use cases requiring more direct control of hardware, Google has introduced C4A Metal, a bare‑metal instance currently available in preview. The service provides native access to Axion processors without hypervisor abstraction, enabling deterministic performance for demanding workloads spanning cloud and edge environments. Potential applications include automotive virtual hardware‑in‑the‑loop testing, native Android build pipelines and specialised enterprise infrastructure.

The Axion portfolio is rounded out by N4A instances, which target cost‑efficient scale‑out workloads such as web services, APIs and data pipelines. Together, C4A, C4A Metal and N4A are intended to provide a unified computing platform that spans AI inference, general cloud workloads and edge deployments, all based on a common Arm architecture.

Arm’s growing role within Google Cloud reflects its broader influence across the cloud and edge computing ecosystem. Google is already running a wide range of production services on Axion processors, including YouTube, Gmail, BigQuery, Spanner and various internal data processing platforms. The company has also migrated tens of thousands of internal applications to Arm‑based infrastructure.

To support customers making a similar transition, Arm and Google Cloud are continuing to expand software ecosystem support, including validated application stacks and migration resources. The companies say this ecosystem maturity is crucial to enabling Arm‑first computing for complex AI workloads as agentic systems become a standard part of enterprise and cloud architectures.

As agentic AI moves towards mainstream adoption, Arm and Google Cloud’s expanded collaboration highlights a shift in how AI infrastructure is being designed – placing greater emphasis on efficient, scalable CPU platforms working alongside specialised accelerators to support increasingly autonomous and complex AI systems.