{"id":153815,"date":"2026-08-27T23:30:12","date_gmt":"2026-08-27T23:30:12","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/153815\/"},"modified":"2026-08-27T23:30:12","modified_gmt":"2026-08-27T23:30:12","slug":"first-benchmarks-revealed-for-openais-jalapeno-ai-chip","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/153815\/","title":{"rendered":"First Benchmarks Revealed for OpenAI&#8217;s Jalape\u00f1o AI Chip"},"content":{"rendered":"<p>\t\t\t\t\t\/\/php echo do_shortcode(&#8216;[responsivevoice_button voice=&#8221;US English Male&#8221; buttontext=&#8221;Listen to Post&#8221;]&#8217;) ?&gt;<\/p>\n<p class=\"wp-block-paragraph\">At a pre-Hot Chips media briefing this week, OpenAI\u2019s VP of hardware, Richard Ho, revealed results of first benchmarks for its clean\u2011sheet ASIC designed specifically for large language models (LLMs). The company\u2019s in\u2011house \u201cJalape\u00f1o\u201d accelerator, developed with Broadcom and Celestica, targets both high throughput and low latency in the same architecture\u2014while addressing what Ho framed as the real enemy in modern AI systems: data movement.<\/p>\n<p class=\"wp-block-paragraph\">OpenAI announced the chip in June, stating it was built from the ground up for current and future LLMs across the industry, developed from design to production in nine months (which itself was accelerated by OpenAI\u2019s models).<\/p>\n<p class=\"wp-block-paragraph\">This week, the company revealed that the tests of the chip and the systems around it showed significant performance advances, with Jalape\u00f1o serving more AI work per unit of power, while also returning responses more quickly. \u201cJalape\u00f1o delivers both higher throughput and lower latency with one architecture, where existing hardware systems often have to make a tradeoff between the two,\u201d OpenAI said in an official statement.<\/p>\n<p><a href=\"https:\/\/www.eetimes.com\/wp-content\/uploads\/image_7249ec.jpeg\" target=\"_blank\" rel=\" noopener nofollow\"><img data-recalc-dims=\"1\" fetchpriority=\"high\" decoding=\"async\" width=\"640\" height=\"360\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/image_7249ec.jpeg\" alt=\"Chip timeline: OpenAI emphasized the speed of development of its Jalape\u00f1o chip. [Image: OpenAI]\" class=\"wp-image-101771646\" style=\"aspect-ratio:1.7777777777777777;width:624px;height:auto\"\/><\/a>Chip timeline: OpenAI emphasized the speed of development of its Jalape\u00f1o chip. (Source: OpenAI)<\/p>\n<p class=\"wp-block-paragraph\">To quantify Jalape\u00f1o\u2019s advantages, Ho said in his briefing that OpenAI leaned on InferenceX, a public benchmark from SemiAnalysis that measures the full serving path for AI requests. On three models\u2014OpenAI\u2019s own GPT-OSS-120B, as well as external models like Kimi K2.5 1T and DeepSeek R1 670B\u2014OpenAI reports that the Jalape\u00f1o chip delivers:<\/p>\n<p>\t\t\t\t\t<a class=\"article-links\" href=\"https:\/\/www.eetimes.com\/reliable-power-path-design-integrating-mosfets-diodes-tvs-devices-and-capacitors\/\" title=\"Reliable Power-Path Design: Integrating MOSFETs, Diodes, TVS Devices, and Capacitors\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" data-recalc-dims=\"1\" loading=\"lazy\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/1787873409_557_Featured-image-24V-DC-industrial-power-path-PCB-circuit-layout-diagram.jpg\" alt=\"Reliable Power-Path Design: Integrating MOSFETs, Diodes, TVS Devices, and Capacitors\"\/><\/a><\/p>\n<p>By Unikey Electronics Pte. Ltd.\u00a0 08.26.2026<\/p>\n<p>\t\t\t\t\t<a class=\"article-links\" href=\"https:\/\/www.eetimes.com\/from-days-to-minutes-accelerating-3d-ic-debug-with-agentic-ai\/\" title=\"From Days to Minutes: Accelerating 3D IC Debug with Agentic AI\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" data-recalc-dims=\"1\" loading=\"lazy\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/1787873409_174_Thumbnail-2-1.jpg\" alt=\"From Days to Minutes: Accelerating 3D IC Debug with Agentic AI\"\/><\/a><\/p>\n<p>By Zackary Glazewski, Founding AI Engineer, ChipAgents\u00a0\u00a0 08.26.2026<\/p>\n<p>\t\t\t\t\t<a class=\"article-links\" href=\"https:\/\/www.eetimes.com\/why-automation-is-essential-to-achieve-eu-cra-compliance\/\" title=\"Why Automation Is Essential to Achieve EU CRA Compliance\u00a0\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" data-recalc-dims=\"1\" loading=\"lazy\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/1787873409_132_BGNetworks-article-thumbnail.jpg\" alt=\"Why Automation Is Essential to Achieve EU CRA Compliance\u00a0\"\/><\/a><\/p>\n<p>By Colin Duggan, CEO and co-founder, BG Networks\u00a0\u00a0 08.26.2026<\/p>\n<p>1.5 to 1.9\u00d7 more AI work per watt at peak throughput versus comparable systems listed on InferenceX.<\/p>\n<p>1.7 to 3.6\u00d7 lower end\u2011to\u2011end latency.<\/p>\n<p>For highly interactive workloads, 2.1 to 4.1\u00d7 higher performance than comparable systems.<\/p>\n<p><a href=\"https:\/\/www.eetimes.com\/wp-content\/uploads\/image_6f4fc1.jpeg\" target=\"_blank\" rel=\" noopener nofollow\"><img loading=\"lazy\" data-recalc-dims=\"1\" decoding=\"async\" width=\"640\" height=\"360\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/image_6f4fc1.jpeg\" alt=\"OpenAI presented several benchmark performance results \u2013 for example this one illustrating peak and matched throughput. [Image: OpenAI]\" class=\"wp-image-101771647\" style=\"aspect-ratio:1.7777777777777777;width:624px;height:auto\"\/><\/a>OpenAI presented several benchmark performance results, for example, this one illustrating peak and matched throughput. (Source: OpenAI)<\/p>\n<p class=\"wp-block-paragraph\">Ho also draws attention to a methodological nuance: OpenAI reports single\u2011token prediction (STP) performance, without speculative multi\u2011token prediction (MTP) decode, while many public numbers include that optimization.<\/p>\n<p class=\"wp-block-paragraph\">\u201cOur numbers are without speculative decode. Something called STP, meaning doing single token prediction, versus some of the numbers you\u2019ll see out there are actually multiple token predictions, which in essence result in about a three to five times boost,\u201d Ho said. \u201cSo, our single token prediction is actually as good or beats some of the multi token prediction of the other models of the other chips.\u201d<\/p>\n<p><a href=\"https:\/\/www.eetimes.com\/wp-content\/uploads\/image_e27446.jpeg\" target=\"_blank\" rel=\" noopener nofollow\"><img loading=\"lazy\" data-recalc-dims=\"1\" decoding=\"async\" width=\"640\" height=\"360\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/image_e27446.jpeg\" alt=\"The pros and cons of speculative decoding using STP vs MTP [Image: OpenAI]\" class=\"wp-image-101771649\" style=\"aspect-ratio:1.7777777777777777;width:624px;height:auto\"\/><\/a>The pros and cons of speculative decoding using STP vs. MTP (Source: OpenAI)<\/p>\n<p class=\"wp-block-paragraph\">When comparing architectures, he said this distinction matters. OpenAI suggests Jalape\u00f1o is competitive even before you turn on aggressive, potentially model\u2011sensitive inference optimizations.<\/p>\n<p>Not just a repurposed GPU; it\u2019s a general-purpose processor for AI workloads<\/p>\n<p class=\"wp-block-paragraph\">During this week\u2019s briefing, Richard Ho emphasized that the new chip is not a repurposed legacy GPU, but a ground-up AI accelerator tailored specifically for transformer workloads. Designed around tight memory-compute co-design, it anchors a multi-generation roadmap. He also reiterated, as in the June announcement, that they\u2019d used AI to accelerate the development of the Jalape\u00f1o chip.<\/p>\n<p><a href=\"https:\/\/www.eetimes.com\/wp-content\/uploads\/image_9ba05b.jpeg\" target=\"_blank\" rel=\" noopener nofollow\"><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" width=\"640\" height=\"360\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/image_9ba05b.jpeg\" alt=\"The Jalape\u00f1o chip is a general-purpose AI accelerator designed for performance\/Watt. [Image: OpenAI]\" class=\"wp-image-101771648\" style=\"aspect-ratio:1.7777777777777777;width:624px;height:auto\"\/><\/a>The Jalape\u00f1o chip is a general-purpose AI accelerator designed for performance\/Watt. (Source: OpenAI)<\/p>\n<p class=\"wp-block-paragraph\">\u00a0\u201cWe basically started with a blank sheet of paper,\u201d Ho said. \u201cWe came in and we looked at the LLM models, where the bottlenecks were, where the inner loops were, and what was going on, and I think we might be the first like large scale chip design team to really do that from scratch without a legacy architecture, without a legacy programming model they had to support.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The result is a general\u2011purpose accelerator for AI workloads, rather than something narrowly tuned to OpenAI\u2019s own models. In the benchmarks, he had emphasized that Jalape\u00f1o ran both internal and external\/open models, highlighting it as a flexible platform rather than a single\u2011model engine.<\/p>\n<p>Data movement is a critical bottleneck<\/p>\n<p class=\"wp-block-paragraph\">Ho added that the defining problem in LLM compute is not just peak TFLOPs, but how much data has to move, and how far.<\/p>\n<p class=\"wp-block-paragraph\">\u201cData movement was super critical,\u201d he said, and became the starting point for the architecture. \u201cIf you\u2019re forced to move that data around to different cores, it just takes up a lot of time and energy. So, we designed an architecture which minimizes that. There is close affinity between HBM banks and cores, and so if you\u2019re able to take advantage of that in your programming model, then you eke out a lot of benefits.\u201d<\/p>\n<p class=\"wp-block-paragraph\">This \u201caffinity\u201d between HBM banks and compute cores is a key design point. Rather than treating the memory system as a flat, uniform pool, Jalapeno binds memory regions closely to compute resources, with the software stack aware enough to exploit that locality.<\/p>\n<p><a href=\"https:\/\/www.eetimes.com\/wp-content\/uploads\/image_0e648d.jpeg\" target=\"_blank\" rel=\" noopener nofollow\"><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" width=\"640\" height=\"360\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/image_0e648d.jpeg\" alt=\"Minimizing data movement is key to optimizing performance. [Image: OpenAI]\" class=\"wp-image-101771650\" style=\"aspect-ratio:1.7777777777777777;width:624px;height:auto\"\/><\/a>Minimizing data movement is key to optimizing performance. (Source: OpenAI)<\/p>\n<p class=\"wp-block-paragraph\">That means model state, including the KV cache used while generating a response, is explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase. The network is integral to the architecture and allows the entire workload to remain within one connected system.<\/p>\n<p class=\"wp-block-paragraph\">According to OpenAI, the result is a balanced, versatile accelerator designed to adapt to shifting model architectures. It excels across both prefill and decode phases and dynamically adjusts to the fluctuating demands characteristic of agentic workloads.<\/p>\n<p><a href=\"https:\/\/www.eetimes.com\/wp-content\/uploads\/image_e92420.png\" target=\"_blank\" rel=\" noopener nofollow\"><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" width=\"640\" height=\"360\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/image_e92420.png\" alt=\"The entire workload remains within one connected system, minimizing data movement and helping the complete request to stay fast and efficient from beginning to end.  [Image: OpenAI]\" class=\"wp-image-101771654\"\/><\/a>The entire workload remains within one connected system, minimizing data movement and helping the complete request to stay fast and efficient from beginning to end.  (Source: OpenAI)<\/p>\n<p>On using AI to build a chip that can be programmed by AI<\/p>\n<p class=\"wp-block-paragraph\">The nine-month tape-out milestone cited by OpenAI was achieved with the use of AI to design the chip itself. AI-driven workflows accelerated implementation exploration, tightened design, measurement, and verification cycles, and enabled continuous iteration across model workloads.<\/p>\n<p class=\"wp-block-paragraph\">In his briefing, Ho said that Jalape\u00f1o is an example of AI assisted hardware design. OpenAI used earlier generation models to assist with chip design and bring up, and is now using its latest models to accelerate software optimization and kernel tuning. According to Ho, AI contributed both to arithmetic circuit optimization and to packing more compute into the die.<\/p>\n<p><a href=\"https:\/\/www.eetimes.com\/wp-content\/uploads\/image_a0725f.jpeg\" target=\"_blank\" rel=\" noopener nofollow\"><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" width=\"640\" height=\"360\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/image_a0725f.jpeg\" alt=\"The nine-month tape-out milestone was achieved using AI to design the chip itself. (Source: OpenAI)\" class=\"wp-image-101771651\" style=\"aspect-ratio:1.7777777777777777;width:624px;height:auto\"\/><\/a>The nine-month tape-out milestone was achieved using AI to design the chip itself. (Source: OpenAI)<\/p>\n<p class=\"wp-block-paragraph\">He noted that the team was able to bring up three separate LLMs on Jalape\u00f1o in about two months\u2014a timeline he cited as evidence of both the architecture\u2019s programmability and the effectiveness of AI assisted workflows.<\/p>\n<p>Getting to roofline performance without thermal throttling<\/p>\n<p class=\"wp-block-paragraph\">In his briefing, Ho talked about balancing compute, memory bandwidth, and networking to get very close to the roofline performance without thermal throttling. He said thermals and power are clearly part of the balancing act, and the Jalape\u00f1o chip targets about 700 W per accelerator, which is aggressive but still within the envelope of modern data center cooling solutions.<\/p>\n<p class=\"wp-block-paragraph\">Ho also stressed that the system is designed to avoid the kind of power\u2011driven throttling engineers have grown accustomed to in high\u2011end GPU deployments.<\/p>\n<p class=\"wp-block-paragraph\">\u201cIt doesn\u2019t thermally throttle. There\u2019s almost no thermal throttling here because it is balanced. It\u2019s coming in at about 700 watts, so it\u2019s well within the range of most cooling solutions.\u201d<\/p>\n<p>Roadmap, deployment<\/p>\n<p class=\"wp-block-paragraph\">Ho re-iterated that Jalape\u00f1o isn\u2019t a one\u2011off project, and that it\u2019s the first in a multi\u2011generation roadmap. He said Gen 2 is already in deep development and Gen 3 is \u201ctaking shape.\u201d Each generation builds on what the team learns and aims to further advance both efficiency and speed.<\/p>\n<p class=\"wp-block-paragraph\">On deployment, OpenAI plans to start small\u2011volume rollouts at the end of this year, with a more significant ramp in 2027. Jalape\u00f1o will coexist with Nvidia hardware and other devices within OpenAI\u2019s fleet. In the near term, Ho expects Jalape\u00f1o to be particularly attractive for low\u2011latency inference, such as ultra\u2011responsive code generation and other interactive workloads, where the combination of reduced data movement, memory affinity, and high bandwidth can be fully exploited.<\/p>\n<p class=\"wp-block-paragraph\">Ho also emphasized that external sales aren\u2019t the focus for the Jalape\u00f1o chip. OpenAI\u2019s own compute demand is growing fast enough that simply meeting internal needs will consume capacity for the foreseeable future.<\/p>\n<p class=\"wp-block-paragraph\">This was also underscored by Sarah Friar, CFO of OpenAI, <a href=\"https:\/\/openai.com\/index\/the-full-stack-behind-abundant-intelligence\/\" rel=\"nofollow noopener\" target=\"_blank\">in a blog<\/a> this week: she said that, while Microsoft\u2019s compute and Nvidia\u2019s chips were foundational to OpenAI\u2019s growth, its portfolio of chips and providers it relies on also includes AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy and SoftBank.<\/p>\n<p class=\"wp-block-paragraph\">\u00a0\u201cWe actively manage this portfolio for both capability and economics. We use premium systems where capability matters most and optimize for efficiency where scale and cost matter more,\u201d she said in the post. \u201cPreserving credible choice across providers, hardware, and deployment models lets us direct demand toward the strongest performance per dollar, maintain pricing discipline as market conditions change, and move with the frontier as stronger technology emerges. Direct control adds leverage where tighter integration can improve the entire system. We partner where the ecosystem helps us move faster and build where co-design creates a meaningful advantage.\u201d<\/p>\n<p class=\"wp-block-paragraph\">This is a balanced way of saying, as with the broader hyperscaler fraternity, they all need so much AI compute to satisfy demand that, in the short to medium term, all options are open and required right now. Hence, they won\u2019t be replacing the GPU incumbents just yet. But with everyone developing their own ASICs and custom chips, there may just come a point in time when the power\/performance tradeoffs push one solution over the tipping point into being the preferred approach.<\/p>\n<p>Read also:<\/p>\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.eetimes.com\/openai-hardware-chief-ai-scaling-laws-will-continue\/\" rel=\"nofollow noopener\" target=\"_blank\">OpenAI Hardware Chief: \u2018AI Scaling Laws Will Continue\u2019<\/a><\/p>\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.eetimes.com\/nxp-expands-industrial-endpoint-access-with-mcu-topology-discovery\/\" rel=\"nofollow noopener\" target=\"_blank\">NXP Expands Industrial Endpoint Access with MCU Topology Discovery<\/a><\/p>\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.eetimes.com\/andes-condor-closure-came-amid-broader-cost-cutting-effort\/\" rel=\"nofollow noopener\" target=\"_blank\">Andes Condor Closure Came Amid Broader Cost-Cutting Effort<\/a><\/p>\n<p class=\"post-tags\">\n<p>        <a href=\"https:\/\/www.eetimes.com\/tag\/ai-and-big-data\/\" rel=\"tag nofollow noopener\" target=\"_blank\">AI AND BIG DATA<\/a>, <a href=\"https:\/\/www.eetimes.com\/tag\/chiplets\/\" rel=\"tag nofollow noopener\" target=\"_blank\">CHIPLETS<\/a>, <a href=\"https:\/\/www.eetimes.com\/tag\/semiconductor\/\" rel=\"tag nofollow noopener\" target=\"_blank\">SEMICONDUCTOR<\/a>, <a href=\"https:\/\/www.eetimes.com\/tag\/soc\/\" rel=\"tag nofollow noopener\" target=\"_blank\">SOC<\/a>\n    <\/p>\n<p class=\"post-tags\">\n<p>        <a href=\"https:\/\/www.eetimes.com\/company\/openai\/\" rel=\"tag nofollow noopener\" target=\"_blank\">OPENAI<\/a>\n    <\/p>\n","protected":false},"excerpt":{"rendered":"\/\/php echo do_shortcode(&#8216;[responsivevoice_button voice=&#8221;US English Male&#8221; buttontext=&#8221;Listen to Post&#8221;]&#8217;) ?&gt; At a pre-Hot Chips media briefing this week,&hellip;\n","protected":false},"author":2,"featured_media":153816,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[35898,64336,157,3123,21767],"class_list":["post-153815","post","type-post","status-publish","format-standard","has-post-thumbnail","category-openai","tag-ai-and-big-data","tag-chiplets","tag-openai","tag-semiconductor","tag-soc"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/153815","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=153815"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/153815\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/153816"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=153815"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=153815"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=153815"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}