null - Seoul Economic Daily International News from South Korea

Google, Amazon, Microsoft (MS) and Meta. These are the hyperscalers that operate massive computing infrastructure. In their early days, their core fields differed — search, distribution, PCs and social media (SNS), respectively — but their business structures are increasingly converging. Only the names of their products and services differ; they are alike in that they make their own artificial intelligence (AI) chips, build cloud services in data centers equipped with those chips, and develop their own AI models.

This kind of business structure is called full stack, or vertical integration. It is a structure that controls every AI layer, from custom chips to the applications (apps) that consumers use. Hyperscalers are pouring their ample cash and enormous investment funds raised from outside into vertical integration. The $700 billion in capital expenditure that major hyperscalers have decided to invest in AI infrastructure this year is also ammunition for building the full stack.

[CAPTIONS]
Conceptual diagram of Google's full stack. Google
Anthropic logo. Reuters/Yonhap
OpenAI CEO Sam Altman (left) and Broadcom CEO Hock Tan hold up jalapeño chip wafers. OpenAI
Google's 8th-generation TPU. Google
AWS logo. AWS
※ Subscribe to Silicon Valley Look for Silicon Valley technology, investment and startup news, along with entertaining reads. https://media.naver.com/hotissue/main?sid1=163&cid=1088902
Maia 200. Captured from MS blog - Seoul Economic Daily International News from South Korea[CAPTIONS]
Conceptual diagram of Google’s full stack. Google
Anthropic logo. Reuters/Yonhap
OpenAI CEO Sam Altman (left) and Broadcom CEO Hock Tan hold up jalapeño chip wafers. OpenAI
Google’s 8th-generation TPU. Google
AWS logo. AWS
※ Subscribe to Silicon Valley Look for Silicon Valley technology, investment and startup news, along with entertaining reads. https://media.naver.com/hotissue/main?sid1=163&cid=1088902
Maia 200. Captured from MS blog

The same is true for the mega-startups that, while not yet matching the hyperscalers, are engaged in fierce AI model competition with them. xAI, the family company of Tesla, along with OpenAI and Anthropic, whose valuations approach $1 trillion, have also jumped into developing their own chips or building data centers. So why are the leading AI players calling for vertical integration?

The pioneer of the full stack is unquestionably Google. Google began its AI business in earnest when it acquired DeepMind, founded by Demis Hassabis, “the father of AlphaGo,” in 2014. The following year, in 2015, it developed its own AI chip, the Tensor Processing Unit (TPU), for the first time, leading the custom AI chip market. In the early stages of development, the TPU was used only internally, but with the arrival of the first-generation Cloud TPU in 2018, companies were able to use the TPU indirectly. In April this year, Google for the first time released TPUs (eighth generation) separated into training and inference versions.

null - Seoul Economic Daily International News from South Korea

Google has built an AI full stack that runs from Stage 1 AI infrastructure, Stage 2 security, Stage 3 research, Stage 4 models and tools, to Stage 5 products and platforms. Google Cloud handles AI chip development, data center construction and security. DeepMind takes charge of research and model development. The culmination of AI technology running from the TPU to Gemini (Google’s AI model) is ultimately implemented in applications that consumers around the world enjoy using, such as Google Search, YouTube and Gmail.

Amazon Web Services (AWS), the world’s No. 1 cloud (virtual storage space connected via the internet) company, has also followed Google in building a full stack. AWS, Amazon’s cloud division, has developed the inference chip “Inferentia” and the training chip “Trainium” since 2018. Although it has not been commercialized like Google’s Gemini, Amazon uses its own AI model, “Nova,” internally.

null - Seoul Economic Daily International News from South Korea

Amazon too has decided to sell its own chips externally as the AI development boom intensified. In a shareholder letter last April, Amazon Chief Executive Officer (CEO) Andy Jassy said, “The estimated annual revenue for our own AI chips has surpassed $20 billion. If we had sold chips this year, annual revenue would have reached $50 billion,” adding, “Because chip demand is very large, there is a very high possibility of selling in bulk to third parties going forward.” This means that while it currently provides chips through leasing AWS data centers, it could sell chips separately.

MS, which trails AWS in second place in cloud, is another Big Tech company that has joined the AI full stack ranks. After unveiling its first AI chip, “Maia 100,” in November 2023, MS released “Maia 200,” with enhanced inference capabilities, in January this year. Maia 200 reflected a situation in which inference performance has become important as the focus of AI shifts from large language models (LLM) to AI agents (assistants) that make their own judgments. MS said that Maia 200’s lightweight computing (FP4) performance reaches three times that of the third-generation version of Amazon’s “Trainium” and surpasses even Google’s seventh-generation Tensor Processing Unit (TPU) “Ironwood.”

Maia has been installed in MS’s data center in Iowa and will also be applied to its facility in Arizona. It is first being deployed for the in-house superintelligence research team’s development and for Copilot, the enterprise AI assistant, and will then be expanded so that customers of its own cloud service, “Azure,” can use it. MS also newly released a software development kit (SDK) that can build AI models based on Maia 200. The SDK is seen as targeting Nvidia’s software platform “CUDA,” which AI developers primarily use.

Meta, a latecomer, has also thrown down the gauntlet to the three cloud companies in order to transform from the operator of Facebook and Instagram into an AI company. Its business model overlaps with those of Google, AWS and MS in that it builds its own system from chip manufacturing to AI model development, differing only in that it does not sell cloud services externally. Recently, reports emerged that Meta sells surplus computing capacity externally, prompting observations that it will enter the cloud business. Meta is investing $50 billion in Louisiana to build a 5-gigawatt (GW) data center, and has followed up by moving to build a facility in Canada as well. Last year, Meta set the goal of developing artificial superintelligence (ASI) that surpasses humans, and after recruiting Alexandr Wang, founder of the AI data startup Scale AI, unveiled its own AI model “Muse Spark.”

Meta, which released its first custom chip in 2018, unveiled four products in the “Meta Training and Inference Accelerator (MTIA)” family — the MTIA 300, 400, 450 and 500 — in March this year, and declared that it would release new chips every six months. MTIA is a custom chip (ASIC) that Meta developed internally to power its data centers. Reuters reported on the 9th, citing an internal memo, that Meta would begin mass production of the new products starting this September. At the March announcement, Giwoon Song, Meta’s Vice President of Engineering, explained, “Releasing a new chip every six months is such a fast cycle that it is very unusual,” adding, “Because we are expanding our current production capacity very quickly and making enormous investments in capital expenditure, we want to be ready to use state-of-the-art chips at any time.”

null - Seoul Economic Daily International News from South Korea

Anthropic and OpenAI, which are competing over the record for the largest initial public offering (IPO), are also striving to build a business structure close to full stack in order to target hyperscale. The two companies, which developed the AI models “Claude” and “GPT” respectively by leasing data centers equipped with Nvidia graphics processing units (GPU) and Amazon and Google chips, are preparing for chip independence.

null - Seoul Economic Daily International News from South Korea

Reuters reported last April, citing sources, that Anthropic was reviewing developing its own chip. The Information reported on the 2nd of this month that Anthropic had begun early-stage work to develop its own AI chip and was in talks with Samsung Electronics’ foundry (contract semiconductor manufacturing), a potential manufacturing partner.

OpenAI, which announced last October that it was cooperating with Broadcom to develop custom chips, unveiled its first inference-only chip “Jalapeño” last month. The Jalapeño chip is expected to be installed, designed according to OpenAI’s vision, in data centers that OpenAI is building.

The biggest reason behind hyperscalers and mega AI startups seeking to build a full stack is to avoid being dependent on giant infrastructure companies such as Nvidia. They regard having the ability to secure AI chips — the core of the AI industry — on their own as their greatest task. Because Nvidia dominates 90% of the data center chip market, the core of AI, AI developers are clinging to securing Nvidia GPUs. They believe that they must have their own chips to escape a structure in which the volume and timing of GPU supply dictate their business.

Another advantage is that having a full stack can boost cost efficiency by building infrastructure optimized for a company’s own AI models. Having their own chips allows them to customize data center design, including the design of cooling and power infrastructure, to obtain more computing power at the same power and in the same footprint. Because GPUs are general-purpose computing chips, [they are ill-suited to] inference models.

null - Seoul Economic Daily International News from South Korea