China’s artificial intelligence startup Moonshot AI has unveiled ‘Kimi K3,’ the world’s largest open-source model with 2.8 trillion parameters, once again sending shockwaves through global financial markets centered on ‘budget-friendly AI.’ With AI semiconductor and infrastructure stocks already in a sharp July sell-off, Wall Street is starkly divided over whether this event will catalyze a rebound or trigger a deeper correction.

The core of the debate hinges on whether intensifying price competition among AI models signals the end of the semiconductor super-cycle, or whether it represents a ‘second launchpad’ poised to ignite even greater demand. The upcoming Q2 earnings and capital expenditure (CAPEX) guidance from Big Tech, beginning with Google parent Alphabet on the 22nd, is expected to serve as a critical inflection point.

China’s Open Model Counteroffensive Armed with ‘Cost-Effectiveness’

Moonshot AI unveiled its next-generation large language model (LLM) ‘Kimi K3’ on the 17th (local time). The model features a context window of 1 million tokens, and the company claims it surpasses OpenAI’s ‘GPT-5.6-Sol’ and Anthropic’s top-tier models on certain coding and AI agent benchmarks in its internal evaluations. However, Moonshot acknowledged that a performance gap remains compared to the most advanced closed-source models, such as Anthropic’s ‘Fable 5.’

What grabbed attention was the disruptive pricing. The API usage fee for Kimi K3 is set at $3 per million input tokens and $15 per million output tokens — less than half the cost of Anthropic’s latest models. The cost per task was analyzed at $0.94, cheaper than GPT-5.6-Sol’s $1.04, half the cost of Opus 4.8 ($1.80), and roughly one-third the cost of Fable 5 ($2.75).

This follows the July 15th launch of ‘Inkling,’ the first model from Thinking Machines Lab, led by former OpenAI Chief Technology Officer Mira Murati. Also released as an open-weight model, Inkling focuses on cost-effectiveness — handling most tasks at a ‘good enough’ level at a much lower price point rather than chasing peak performance.

With these moves, the price war within the open model camp is intensifying, now including US-based Thinking Machines Lab alongside China’s DeepSeek and GLM, following the ‘DeepSeek shock’ of January 2025. If the average price of AI models continues its downward trajectory, margin compression for model vendors appears inevitable.

Peak Cycle Debate for Semiconductors Reignites

This trend is fueling fundamental market skepticism about AI infrastructure investment. The Philadelphia Semiconductor Index, which surged 89% in Q2 alone, has dropped 15% in July, while the memory chip ETF ‘DRAM’ has plunged over 20%. The correction in AI hardware stocks — spanning power, optical communications, and data center infrastructure — is also proving prolonged.

Investors are increasingly anxious about whether the astronomical capital poured into infrastructure by frontier model developers like OpenAI, Anthropic, and Google can generate sufficient return on investment (ROI). There are concerns that if low-cost models from China proliferate, the pace of massive CAPEX deployment by US Big Tech could decelerate. Meta is currently investing over $100 billion in AI infrastructure this year, while OpenAI is pursuing the ‘Stargate’ project in partnership with SoftBank and Oracle.

Conversely, Big Tech software companies previously weighed down by facilities investment burdens have staged a relative rebound. This reflects growing expectations that economic value within the AI value chain will diffuse from ‘physical bottlenecks’ like GPUs and HBM toward revenue-generating applications.

Bulls Bet on the ‘Jevons Paradox’

However, prominent tech bulls like Gavin Baker, Chief Investment Officer at Atreides Management, interpret this as the beginning of an ‘AI infrastructure super-bull scenario.’ The logic rests on the ‘Jevons paradox’ — the notion that increased efficiency and lower costs paradoxically drive an explosion in usage.

Baker estimates that Anthropic’s and OpenAI’s inference margins currently approach 90%. He argues that as open models proliferate and these margins are transferred to consumers, the universe of tasks amenable to AI adoption will expand explosively, ultimately driving greater total token consumption and computing demand. NVIDIA’s strategy of directly releasing open-source models and supporting open ecosystems is seen as a move to diversify demand — currently concentrated among a handful of frontier model developers — toward mainstream enterprises, governments, and research institutions.

Ultimately, the current shift likely represents not a ‘migration’ of AI market value from semiconductors to applications, but rather a ‘diffusion’ across the entire value chain. However, the market may have already priced in these expectations too aggressively.

Eyes on Big Tech Earnings Starting July 22

The key question is whether usage growth is outpacing the rate of AI price declines, and whether hyperscalers can grow revenue faster than costs to sustain their CAPEX expansion ambitions. Consequently, Q2 earnings and CAPEX guidance from major Big Tech players — starting with Google (Alphabet) on the 22nd, followed by Microsoft and Amazon — have emerged as the pivotal variables that will dictate semiconductor investment sentiment.

Moonshot AI, meanwhile, is a startup founded in Beijing in 2023 by CEO Yang Zhilin, a former Tsinghua University professor. It has attracted large-scale investments from Alibaba and HSG, among others, pushing its valuation to approximately $31.5 billion. Annual recurring revenue (ARR) has surpassed $200 million, and the company is also pursuing a listing on the Hong Kong stock exchange.