Everybody wants tokens, but they also want amazing economics.

So says Chris Lattner, co-Founder & CEO of Modular, pitched as the portable alternative to NVIDIA’s software stack, designed for every AI accelerator, and soon to be part of the Qualcomm empire following a $3.9 billion takeover offer.

Compute has fundamentally changed, argues Lattner:

It’s no longer about a single chip. Compute today is a large-scale data center distributed systems problem. We all need to program diverse AI accelerators from multiple different vendors. We need to get the best performance, the best TCO, and we need usability because doing all this is harder than it’s ever been before. Now the world is still struggling to get individual systems to compete with the industry leader, but that’s where Modular comes in. 

At Modular, we spent the last 4.5 years building a novel platform that actually scales, all with the goal from the beginning of unifying the industry and opening a new chapter for accelerated compute. Now we built this platform to scale across the full spectrum, starting from the data center, but then going all the way down to the Edge.

The end result, according to his co-Founder Tim Davis is a full AI compute platform for the heterogeneous data center:

Importantly, our platform turns heterogeneous data center systems into multi-silicon AI token factories. In this world, enterprises, partners and developers can use the best silicon for each workload without being locked into a single hardware stack. Because Modular is heterogeneous by design, the industry can achieve lower TCO, higher performance and greater portability across the world’s compute infrastructure.

Token amounts

And that obviously is a big deal as tokenomics – aka soaring costs – come into play for end users. The appeal to a firm like Qualcomm is obvious as CEO Christiano Amon notes:

Agents generate demand for a lot of tokens. The reason a lot of the hyperscalers just see a wall in front of them of compute demand is because the economics of AI are fundamentally changing…Agents and orchestrators, they’re re-defining the architecture and economics of AI, not only creating an entry point for Qualcomm into the data center, but actually, creating a fundamental change in the architecture of compute that touches all the devices on the edge.

If you look at how we started with conversational to now agents, you see the order of magnitude increase, the projection is 40x the increase in annual token demand between 2026 and 2030…You can actually see how the architecture of AI is evolving.

It’s important not to get distracted at this point, he argues:

You can see lot of useless debates about is this edge or this is cloud? It is actually the wrong conversation to say, ‘I have something that need to do on the cloud. Can I do it on the edge ?’, and vice versa. That’s the wrong approach. Things that are going to be done in the cloud are going to be done in the cloud. The growth of the cloud is incredible, but the edge now also become a computing [platform] that is going to generate tokens and I think how the industry is naturally going to evolve.  What’s happened in the data center [is that] inference is actually becoming dis-aggregated and distributed everywhere. This is a new form of compute.

Economical thinking

So if agentic AI does indeed change the economics of compute, what does that mean in practice? Antonios Pialis, General Manager of Data Center for Qualcomm, posits:

It means token counts are skyrocketing as we introduce agents. It means CPU attach rates are soaring through the roof. You can’t find CPUs anywhere. They’ve already been bought up. So traditional infrastructure will not scale to the needs of agentic AI. So the industry needs a paradigm shift in order to deliver this.

Agentic AI requires a new compute infrastructure for a good reason, he adds:

The industry has blown through gen AI and reasoning, and here we are in the cusp of deploying agents en masse. But what does that mean? It means a single agent’s queries generating 50x to 100x inference requests. We have over 1 million tokens being generated by a single query. Traditional compute infrastructure cannot support the scale.  A paradigm shift is needed.

The traditional GPU-based compute that’s been deployed across data centers worldwide was built to support both training and inference and runs hundreds of kilowatts today, he goes on, while solutions will be running north of 500 kilowatts. How can enterprises deploy this and support 100x more inference calls than they do today? Pialis’ answer is blunt:

We can’t!

New solutions

So a new solution is needed and that, as far as Qualcomm is concerned, takes the form of its new dis-aggregated compute infrastructure. Pialis explains:

What’s needed to lead in data center [is] bespoke solutions that deliver hardware acceleration for each and every function needed to deploy agentic compute [with] various forms of CPUs performing specific functions, unique XPUs, some targeted to attention for pre-fill, some targeted for KV (Key Value) caching during de-code. 

Transformer sizes are growing 240x over a span of two years these days and compute memory is only doubling in that same time span, he observes. This has further implications:

It means there’s no point packing more compute unless we solve the memory bottleneck. We have broken through the memory bottleneck. How have we done it? We have rearchitected compute for XPUs. We have separated the AI accelerator from the XPU. And what you see, we now put our XPU right under a DRAM (Dynamic Random Access Memory) stack. What does this mean? We offer all of the performance advantage of SRAM (Static Random Access Memory), but with the density and the memory capacity that HBM (High Bandwidth Memory) stacks offer.

That means that congestion that is seen with HBM is gone, he adds:

A great analogy is imagine working in the same building that you live in. And so you only travel up and down. And what does that mean for the highways and the roads that connect the suburbs to the city? Guess what? The roads are clear. So the value this brings to the industry is lower power consumption, less heat. And that expensive road of silicon interposer that HBM solutions use are no longer needed.

We can deploy multiple HBC (High Bandwidth Compute) stack within a single compute device using standard packaging. That is a tremendous value that we deliver to the industry in terms of performance per cost advantage..HBC offers 200x capacity per watt, better solution than SRAM…With HBC, we deliver 6x the bandwidth per watt versus competitor HBM-based solution…With HBC, we offer a single solution that can seamlessly span this entire sphere of workload and deliver multiple fold performance per watt and performance per dollar benefit. That is a direct TCO advantage that we offer to the industry.

And on top of the hardware stack is software, he goes on:

Software is where the magic is, and it takes a lot for me, a hardware designer, to acknowledge that. We will be deploying a full software solution stack that includes the most sophisticated orchestrators that will manage and route the traffic across a disaggregated compute cluster, all the way down through frameworks. And most importantly, open frameworks that will allow model developers to both develop and deploy at scale their models. While others in the industry build moats trying to protect their hardware deployments, we at Qualcomm believe in building bridges to unite the industry.

We have developed a transformational infrastructure that is already winning in the industry. Four product lines, each of them already anchored with multiple customer wins and a pipeline that will blow your heads in terms of accumulated value. Incredible metrics. Look at that, up to 8x better tokens per watt per second than traditional GPUs, greater than 200x memory capacity compared to SRAM solutions, 6x memory bandwidth per watt. And for our CPUs, greater performance than 2x than our competition. All of this is direct TCO advantage.

He concludes:

We have the performance that the industry needs. Tokens-per-watt replaces FLOPS (Floating Point Operations Per Second). The race has changed. Embedded solutions, embedded providers, they’re playing the old game. There’s a new game in town, and it’s all about delivering agentic first rack scale platforms that delivers the world’s best TCO.

My take

A bold new approach. No flops ahead? Will it add up for end users? Time will tell.