Samsung Electronics has unveiled “zHBM,” a next-generation architecture that vertically stacks High Bandwidth Memory (HBM) directly on top of AI accelerators, targeting the coming era of AI agents. The goal is to boost AI system response speeds tenfold—from the current 100 tokens per second to 1,000 tokens per second.

Kim In-dong, senior vice president of memory product planning at Samsung Electronics’ Device Solutions Americas (DSA), presented the vision during a keynote address at the AI Infrastructure Summit held on the 16th (local time) at the Santa Clara Convention Center in California. “Current conversational AI systems remain at roughly 100 tokens per second per user,” Kim said. “To prepare for the coming agentic AI era, we are targeting a quantum jump to 1,000 tokens per second—a tenfold improvement.”

3D Stacking Architecture That Transcends 2.5D Limitations

The industry currently relies primarily on 2.5D packaging, which places HBM side-by-side with AI accelerators on a planar surface. However, this approach has been criticized for structural limitations—long data transfer distances and constrained pathways—that create latency bottlenecks.

Samsung Electronics’ zHBM addresses this problem through a 3D structure that stacks HBM vertically, directly on top of the accelerator. Kim likened the architecture to “installing a dedicated express elevator that goes straight from a hotel room to the first-floor lobby,” explaining that it can dramatically reduce data travel distance and bottleneck-induced latency.

According to Samsung Electronics, zHBM can deliver up to 8x the performance and more than 3x the power efficiency of HBM5.

MetricConventional 2.5D PackagingzHBM (3D Stacking)HBM PlacementPlanar, beside acceleratorVertical, atop acceleratorPerformance (vs. HBM5)BaselineUp to 8xPower Efficiency (vs. HBM5)Baseline3x or greaterData Transfer DistanceRelatively longDramatically shortened

Note: Performance and power efficiency figures are targets relative to HBM5, as presented by Samsung Electronics at the AI Infrastructure Summit.

Regarding concerns that heat generated by the accelerator could transfer to the memory in a vertical stacking configuration, the company plans to address this through a co-design system that involves customized collaboration with AI accelerator customers from the earliest stages of product design.

zNAND-O Roadmap Targets On-Device AI

Samsung Electronics also unveiled its roadmap for “zNAND-O,” a NAND flash-based 3D storage solution targeting the on-device AI market.

“By 2030, it will become commonplace for trillion-parameter AI models to run locally in environments such as AI workstations, rather than on large-scale servers,” Kim said. “Building memory for a trillion-parameter model using DRAM alone would be prohibitively expensive, but zNAND-O can achieve this at one-sixth the cost.”

Samsung Electronics plans to begin customer sampling of zNAND-O in earnest starting in 2028. “We will provide unmatched value to customers in ‘time-to-market,’ which is the core of AI semiconductor competition,” Kim added.

Significance of the Memory Paradigm Shift

The announcement is being interpreted as a strategy to resolve memory bandwidth and latency issues—long identified as the bottleneck in AI computing—through a fundamental change in packaging architecture. Given that AI agents performing complex multi-step tasks beyond simple Q&A will require dramatic improvements in token generation speed, the 1,000-tokens-per-second target set by zHBM is expected to become a key benchmark in the next-generation AI infrastructure race.

In the on-device space, zNAND-O is seen as a card that could accelerate the democratization of AI workstations—addressing growing demand to run massive AI models without large-scale servers while cutting costs to one-sixth of conventional DRAM-centric configurations.