ICS, AMS, and CCE VolcanoNext form the backbone of Huawei’s bid to run AI agents at massive scale on homegrown silicon

Huawei new office in Jakarta, Indonesia

– Getty Images

Huawei Cloud has unveiled an “Agentic Infra” stack: a full suite of compute, storage, and networking products designed to run AI agents at massive scale on its NPU-based cloud.

What looks to be the cloud vendor’s most direct bid yet to compete with Nvidia-anchored AI infra, Huawei’s Inspire event in Shanghai saw the launch of AICS (AI Cluster Service), a compute platform it claims can support clusters containing 100,000 cards.

Running on Huawei’s own UnifiedBus (UB) interconnect protocol, AICS can provide throughput of five million tokens per second across 1,000 cards – for total computing power of up to a staggering 200 exascale floating-point operations per second (EFLOPS) – all with sub-10 millisecond token-generation latency.

Also unveiled was a storage solution dubbed AMS (Agentic Memory Storage), which Huawei claims can provide memory expansion for its NPU chips, while also tiering key value (KV) cache to further reduce inference costs for long-running agent tasks.

Rounding out Huawei’s Agentic Infra stack is CCE VolcanoNext, a scheduler claiming >30% better resource utilization by pooling training and inference workloads together rather than siloing them, and AgentSphere, a security-isolated sandbox environment where users can spin up hundreds of thousands of agent instances per minute.

Huawei headquarter office building in Vilnius

26 May 2026

Huawei’s τ Scaling Law argues the future of chips lies in time, not transistor size

The stack was formally unveiled during a keynote by Dr. Peter Zhou, director of the Board at Huawei and CEO of Huawei Cloud, who said that agentic AI was driving a fundamental shift in computing paradigms.

Huawei’s stacked Infra showcase at Inspire comes as China pushes to build domestic alternatives, with the giant doubling down on its computing prowess to take advantage of the dearth of American chips in the country following U.S. import bans.

While Huawei CEO Ren Zhengfei admitted last summer that its chips were one generation behind those made by its U.S. counterparts, the company is looking to close the gap quite rapidly. Driving that change is Tau (τ), its own scaling principle for semiconductor design that focuses on improving chip designs without being able to shrink transistors further by reducing propagation delays in chip signals – like by shortening wiring paths to allow for higher performance.

Already, Huawei has used the concept to design some 381 chips, and is set to be combined with its LogicFolding architecture, which already boosts τ performance across the device level, circuit level, chip level, and system level. The latter technology instrumental in developing its hotly anticipated Kirin processor line.

The Inspire-unveiled Agentic Infra stack then promises to let enterprises build and run AI agents at scale on homegrown chips, while utilizing cluster-spanning tools and security solutions that would allow it to go toe-to-toe with many of its American rivals.

A company statement sees Huawei Cloud pledge to drive “software-hardware synergistic innovation to build the foundation for enterprise-grade AI innovation, working alongside global customers, partners, and developers to usher in a brand-new era of agentic AI.”

Dr. Peter Zhou, Director of the Board at Huawei and CEO of Huawei Cloud, delivering the keynote speech at Inspire 2026

Dr. Peter Zhou delivering the keynote speech at Inspire 2026

– Huawei

Models, agents, and the rest of the stack

Beyond the infra stack itself, Huawei also used Inspire to push out a model platform dubbed ModelArtsNext, which adds reinforcement-learning-as-a-service (RLaaS) and a model-routing layer that dynamically sends requests to whichever of its more than 20 partner models, including systems from DeepSeek, Zhipu AI, and MiniMax, is best suited to the task.

Huawei claims the routing engine hits a 95%+ scheduling accuracy and cuts inference costs by around 20%, framing it as a way for enterprises to tap state-of-the-art models without committing to a single provider.

The partner roster was formalized into what the firm labeled an “AI Model Partner Program,” with Huawei positioning its cloud platform as a neutral hosting and routing layer across China’s increasingly crowded domestic model ecosystem.

On the agent-building side, Huawei released AgentArts, an enterprise agent platform aimed at production-grade, long-running agentic tasks, alongside an open-source counterpart, sharing more than 90% of its codebase with the commercial version, and a new portal called AgentArts Orchard for building and deploying agents via command-line interfaces (CLI).

The company also rolled out a dedicated security layer for the stack, including hold-your-own-key (HYOK) hardware encryption and confidential-computing support across virtual machines, training, and inference, while touting more than 1,000 days without a major service incident.

More in IT Hardware & Semiconductors