{"id":151393,"date":"2026-08-26T02:10:20","date_gmt":"2026-08-26T02:10:20","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/151393\/"},"modified":"2026-08-26T02:10:20","modified_gmt":"2026-08-26T02:10:20","slug":"openais-first-custom-chip-jalapeno-reportedly-beats-nvidias-blackwell-and-rubin-in-inference-benchmarks","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/151393\/","title":{"rendered":"OpenAI&#8217;s first custom chip &#8220;Jalape\u00f1o&#8221; reportedly beats Nvidia&#8217;s Blackwell and Rubin in inference benchmarks"},"content":{"rendered":"<p>OpenAI showed off the first benchmarks for its in-house inference chip at the Hot Chips conference. &#8220;Jalape\u00f1o&#8221; reportedly outperforms both Nvidia&#8217;s Blackwell and Rubin in throughput per watt and token latency.<\/p>\n<p>The chip handles inference only, meaning it runs AI models but doesn&#8217;t train them. Jalape\u00f1o isn&#8217;t tuned to OpenAI&#8217;s own models either. It&#8217;s a general-purpose LLM inference accelerator.<\/p>\n<p>OpenAI claims Jalape\u00f1o delivers 1.5x to 1.9x more AI work per watt at peak throughput across all three tested models, with 1.7x to 3.6x lower end-to-end latency than the best commercially available systems. For interactive workloads, the company says performance is 2.1x to 4.1x higher.<\/p>\n<p>The results come from tests using <a href=\"https:\/\/newsletter.semianalysis.com\/p\/openai-jalapeno-better-than-nvidia\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">SemiAnalysis&#8217;s public InferenceX benchmark<\/a>. OpenAI provided the numbers. SemiAnalysis verified some runs on-site in the lab. The models tested were GPT-OSS 120B, Deepseek R1 670B, and Kimi K2.5 1T. On GPT-OSS, Jalape\u00f1o hit about 1,400 tokens per second per user. On Deepseek R1, it topped 700 tokens per second on a single concurrent request.<\/p>\n<p><img fetchpriority=\"high\" decoding=\"async\" class=\"wp-image-39427 size-full\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/openai_chip_benchmark.png\" alt=\"Jalape\u00f1o erzielt bei gleicher Decoding-Geschwindigkeit je nach Modell das 54- bis 104-Fache des Token-Durchsatzes pro Kilowatt im Vergleich zum besten verf\u00fcgbaren Beschleuniger. | Bild: OpenAI\" width=\"1664\" height=\"1116\"\/>At matched decoding speed, Jalape\u00f1o achieves 54x to 104x the token throughput per kilowatt compared to the best available accelerator, depending on the model. | Image: OpenAI<\/p>\n<p>Jalape\u00f1o posted these numbers without using techniques like <a href=\"https:\/\/the-decoder.com\/google-speeds-up-gemma-4-threefold-with-multi-token-prediction\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">multi-token prediction<\/a> or <a href=\"https:\/\/the-decoder.com\/deepseeks-dspark-boosts-ai-speed-by-up-to-85-percent-a-strategic-win-under-tightening-us-export-controls\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">speculative decoding<\/a>, while some of the comparison systems did rely on those optimizations, so there&#8217;s still room to improve.<\/p>\n<p>In its headline performance-per-watt comparison, &#8220;Jalape\u00f1o smokes every other chip,&#8221; <a href=\"https:\/\/newsletter.semianalysis.com\/p\/openai-jalapeno-better-than-nvidia\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">SemiAnalysis writes<\/a>. <a target=\"_blank\" rel=\"noopener nofollow\" href=\"https:\/\/x.com\/dylan522p\/status\/2092258594628706778\">SemiAnalysis CEO Dylan Patel added<\/a>,\u00a0&#8220;Usually first generation chips aren&#8217;t competitive, but OpenAI is beating Nvidia Blackwell and even Rubin.&#8221;<\/p>\n<p>SemiAnalysis points out that the fairer comparison isn&#8217;t Blackwell but Nvidia&#8217;s newer <a href=\"https:\/\/the-decoder.com\/ces-2026-nvidia-promises-five-times-the-ai-performance-and-ten-times-cheaper-inference-with-vera-rubin\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Vera Rubin platform<\/a>, since both use HBM4 memory. Even here, Jalape\u00f1o squeezes out more output tokens per megawatt than Vera Rubin, even though Nvidia&#8217;s accelerator uses the multi-token prediction optimization that Jalape\u00f1o hasn&#8217;t adopted yet. On total cost of ownership per token, the two come out roughly even.<\/p>\n<p>There are caveats, though. Nvidia and AMD have already published results with larger models like Deepseek V4 Pro and Kimi K3 that haven&#8217;t been tested on Jalape\u00f1o yet. And while Rubin systems are already shipping to customers, Jalape\u00f1o reportedly hasn&#8217;t moved beyond engineering samples.<\/p>\n<p>OpenAI built the chip in nine months, partly using its own models<\/p>\n<p>OpenAI developed Jalape\u00f1o with Broadcom. Design work kicked off in mid-2024, and the final design went to fabrication in November 2025. The full cycle took about 16 months, but OpenAI says only nine months passed between the first chip design and the finished blueprint heading to the factory. The company used its own AI models during development, <a href=\"https:\/\/openai.com\/index\/jalapeno-first-results\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">according to OpenAI<\/a>. Older model generations helped with chip design, while newer ones sped up programming and optimization.<\/p>\n<p>SemiAnalysis sees this as a sign that Nvidia&#8217;s much-discussed <a href=\"https:\/\/the-decoder.com\/amds-software-woes-leave-nvidia-unchallenged-in-ai-chip-market-study-finds\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">&#8220;CUDA moat&#8221;<\/a> may not hold anymore. &#8220;The CUDA moat is potentially dead given how fast OpenAI can bring up new models on their silicon,&#8221; the firm wrote.<\/p>\n<p>OpenAI CFO Sarah Friar says the chip fits into a <a href=\"https:\/\/openai.com\/index\/the-full-stack-behind-abundant-intelligence\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">broader compute strategy<\/a> where data centers, chips, models, the developer platform, products, and devices all work as one integrated system. She claims Jalape\u00f1o complements OpenAI&#8217;s existing partnerships with Nvidia, AMD, AWS, Cerebras, CoreWeave, and others rather than replacing them.<\/p>\n<p>OpenAI has deep ties with several of these companies. <a href=\"https:\/\/the-decoder.com\/nvidia-reportedly-set-to-invest-30-billion-in-openai\/\" rel=\"nofollow noopener\" target=\"_blank\">Nvidia<\/a>, <a href=\"https:\/\/the-decoder.com\/amd-signs-a-long-term-deal-to-supply-openai-with-multiple-generations-of-instinct-gpus\/\" rel=\"nofollow noopener\" target=\"_blank\">AMD<\/a>, and <a href=\"https:\/\/the-decoder.com\/openais-growth-obsession-continues-as-aws-becomes-the-latest-megadeal-in-its-scaling-spree\/\" rel=\"nofollow noopener\" target=\"_blank\">AWS<\/a> are all investors or compute partners, with Nvidia being one of the largest. Each of them is also building its own AI chips, which makes the relationship both cooperative and competitive. That said, all of these companies <a href=\"https:\/\/the-decoder.com\/sam-altman-says-scaling-up-compute-is-the-literal-key-to-openais-revenue-growth\/\" rel=\"nofollow noopener\" target=\"_blank\">keep saying the world can&#8217;t have enough compute<\/a>, a claim that conveniently supports their own business models.<\/p>\n<p>\t\t\t\tAI News Without the Hype \u2013 Curated by Humans<\/p>\n<p>\n\t\t\t\t\tSubscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive &#8220;AI Radar&#8221; frontier report six times a year, full archive access, and access to our comment section.\t\t\t\t<\/p>\n<p>\t\t\t\t<a href=\"https:\/\/the-decoder.com\/subscription\/\" class=\"inline-block text-white bg-(--heise-primary) mt-3 hover:bg-blue-800 focus:ring-4 focus:outline-none focus:ring-blue-300 font-medium rounded-sm w-full sm:w-auto  pl-3 pr-3 py-2.5 text-center newsletter-submit-button hover:no-underline\" rel=\"nofollow noopener\" target=\"_blank\"><br \/>\n\t\t\t\t\tSubscribe now\t\t\t\t<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"OpenAI showed off the first benchmarks for its in-house inference chip at the Hot Chips conference. &#8220;Jalape\u00f1o&#8221; reportedly&hellip;\n","protected":false},"author":2,"featured_media":151394,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[4518,2306,157],"class_list":["post-151393","post","type-post","status-publish","format-standard","has-post-thumbnail","category-openai","tag-ai-chips","tag-chips","tag-openai"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/151393","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=151393"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/151393\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/151394"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=151393"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=151393"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=151393"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}