
OpenAI logo displayed on a digital screen in San Francisco, California, August 20, 2026. (Photo by Smith Collection/Gado/Getty Images)
Gado via Getty Images
A processor that will not ship in any volume until 2027 was benchmarked in public this week, against a competitor’s shipping product, on a scoreboard maintained by somebody else. OpenAI put out the first performance figures for Jalapeño, the inference accelerator it designed with Broadcom, and the numbers were strong enough to carry the week’s headlines.
Semiconductor companies publish comparative benchmarks because they are trying to sell silicon. OpenAI has no silicon to sell. It has roughly a decade of data centre construction to pay for, and a public efficiency number speaks directly to the people underwriting it.
The tests ran on SemiAnalysis’s InferenceX benchmark, a public framework that scores the full path of serving a request rather than a synthetic kernel. A 700-watt Jalapeño part delivered up to 1.9 times the throughput per kilowatt of a 1,400-watt Nvidia flagship, with latency as much as 3.6 times lower. Testing ran across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 at a trillion parameters. On DeepSeek R1 with a single user attached, the chip cleared 700 tokens per second.
What The Fine Print Says
SemiAnalysis published its own read alongside the results, and it is unusually direct about the limits of the comparison.
“We believe that comparison to Blackwell is somewhat incomplete and unfair. Jalapeño is really competing against chips like Rubin that also use HBM4.”
Jalapeño carries HBM4 memory at 15.4 terabytes per second of bandwidth. The Blackwell systems it was measured against carry the previous generation. A meaningful share of the efficiency gap is a memory generation showing up as a chip result, and Nvidia’s Rubin parts, which use the same memory class, are the like-for-like comparison nobody has run yet.
Three other limits matter. The silicon tested was A0 stepping, the first working version, with a B0 revision already in fabrication. The workloads were short-context and single-turn, with none of the multi-turn agent traffic that stresses routers and cache behaviour in production. And the numbers themselves came from OpenAI, with SemiAnalysis witnessing the runs in person rather than executing the suite independently.
Those limits are what make the result readable. A first-silicon part hitting those figures on a public benchmark is a real engineering outcome, and the caveats describe precisely which second result would settle the question.
Why Publish It Now
Richard Ho, who runs OpenAI’s hardware program, said Jalapeño arrives “at the end of 2026 in very small volumes” with meaningful deployment in 2027. Publishing detailed competitive numbers more than a year ahead of that is not how product launches are usually sequenced.
The answer sits in how this buildout is being paid for. Nvidia agreed this month to backstop up to $105 billion of financing for an OpenAI data centre in Ohio. That covers an initial 4.25 gigawatts with an option on 3.75 more, and the capacity arrives in phases from 2028. Anthropic signed a $45 billion, six-year compute agreement with Nscale in the same window.
These are project finance structures, sized like toll roads and power plants, and they are being arranged years before the revenue that services them exists. Money on that scale gets committed against a cost forecast, and somebody has to believe the forecast.
Infrastructure capital prices on unit economics. A road gets underwritten on traffic per lane, a plant on output per megawatt-hour, and an inference business on how much useful work comes out of each watt going in. A published performance-per-watt figure is the closest thing OpenAI has to a demonstrated cost curve, and it arrived in the same month the company was arranging nine figures of construction credit per gigawatt.
The choice of venue fits the same logic. Running on an independent public benchmark, with a third party in the room, produces a number that survives diligence in a way a company slide does not.
Broadcom Is The Other Half
Jalapeño exists because of a partnership announced in October 2025. OpenAI and Broadcom committed to deploy 10 gigawatts of OpenAI-designed accelerators, with racks starting in the second half of 2026 and the program completing by the end of 2029. Ten gigawatts is close to a decade of construction, and it is already under contract.
OpenAI did the architecture. Broadcom does the physical implementation, the networking and the manufacturing relationship, which is the work that turns a design into a rack that runs. Every efficiency claim in this week’s benchmark is also a demonstration that the custom accelerator path Broadcom sells to its largest customers produces competitive parts on first silicon.
The companies with a defensible position in this cycle are the ones that get paid whichever way the buyer decides to build. A design partner charging for the hard part of custom silicon sits on the same side of that trade as the merchant vendor selling complete systems, which is why a strong Jalapeño result reads as confirmation for both of them.
What Would Settle It
Two results would turn this from a strong first showing into something structural, and both are specific enough to write down.
The first is a Jalapeño number against Rubin, on parts using the same memory generation. That comparison removes the largest confound in the current data and tells you whether the advantage lives in the architecture or in the procurement calendar.
The second is a multi-turn agent workload. Short-context single-turn traffic is the friendliest case for a specialised design, and production traffic for a company running consumer chat and coding agents looks nothing like it.
There is also a simpler tell available every quarter. OpenAI has continued to commit to Nvidia systems while building its own, which is the same pattern Google and Apple established years ago while running elite internal silicon programs. Should OpenAI’s orders for merchant hardware keep growing alongside the Jalapeño ramp, the honest conclusion is that custom accelerators are adding capacity rather than replacing it. That is the outcome the current evidence points to, and it is the one the financing structures assume.