{"id":151372,"date":"2026-08-26T01:38:13","date_gmt":"2026-08-26T01:38:13","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/151372\/"},"modified":"2026-08-26T01:38:13","modified_gmt":"2026-08-26T01:38:13","slug":"openai-chip-beats-nvidia-systems-with-1-9x-more-work-per-watt","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/151372\/","title":{"rendered":"OpenAI chip beats Nvidia systems with 1.9x more work per watt"},"content":{"rendered":"<p class=\"wp-block-paragraph\">OpenAI says its first custom inference chip, Jalape\u00f1o, can deliver more AI work per watt while also reducing response times, pointing to a hardware design aimed at handling increasingly demanding model workloads more efficiently.<\/p>\n<p class=\"wp-block-paragraph\">The company tested Jalape\u00f1o across three public models: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. Across the tests, the chip delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems.<\/p>\n<p class=\"wp-block-paragraph\">For highly interactive workloads, OpenAI said Jalape\u00f1o achieved 2.1 to 4.1 times higher performance. The company argues that combining throughput and low latency in one architecture could help reduce the hardware and power needed to serve AI models.<\/p>\n<p class=\"wp-block-paragraph\">The chip is rated at 700 watts, although OpenAI said its measured sustained power remained at or below 550 watts during the tested workloads. The company compared Jalape\u00f1o with commercially available accelerator systems using the public InferenceX benchmark from SemiAnalysis.<\/p>\n<p>One chip tackles both bottlenecks<\/p>\n<p class=\"wp-block-paragraph\">A major focus of the design is avoiding the usual tradeoff between throughput and latency. AI inference has different demands depending on what the system is doing. Processing a user\u2019s prompt, known as prefill, is generally compute-intensive, while generating the response token by token, known as decode, depends more heavily on memory bandwidth.<\/p>\n<p class=\"wp-block-paragraph\">Jalape\u00f1o was designed to handle both phases within the same architecture. OpenAI said the chip keeps model state, including the key-value cache used during generation, closer to the processing resources that need it.<\/p>\n<p class=\"wp-block-paragraph\">Its networking system is also integrated into the architecture to reduce the amount of data that needs to move between chips. That is intended to limit communication delays that can leave computing resources waiting for data.<\/p>\n<p class=\"wp-block-paragraph\">The approach becomes particularly relevant for AI agents, which may perform many inference steps in sequence. A small delay in each step can add up to a much longer overall task.<\/p>\n<p>AI helped build Jalape\u00f1o<\/p>\n<p class=\"wp-block-paragraph\">OpenAI also used its own models during the chip\u2019s development. The company said AI helped engineers explore implementations, shorten design and verification cycles, and optimize arithmetic circuits.<\/p>\n<p class=\"wp-block-paragraph\">The team moved from initial design to tapeout in nine months. OpenAI then used Codex with GPT-Astra to bring three <a href=\"https:\/\/interestingengineering.com\/ai-robotics\/nvidia-open-weight-models-china-us\" target=\"_blank\" rel=\"dofollow noopener\">open-weight models<\/a> that were not part of the original production plan to high performance on Jalape\u00f1o within two months.<\/p>\n<p class=\"wp-block-paragraph\">For selected GPT-OSS attention and mixture-of-experts blocks, AI-generated implementations ran 1.5 to 1.8 times faster than existing implementations written by human experts. OpenAI stressed that these results applied to individual blocks rather than the complete models.<\/p>\n<p class=\"wp-block-paragraph\">The company also tested Jalape\u00f1o against large models including Kimi K2.5 1T. On that model, it reported about 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than the comparison system.<\/p>\n<p class=\"wp-block-paragraph\">OpenAI plans to begin deploying Jalape\u00f1o in its compute <a href=\"https:\/\/interestingengineering.com\/ai-robotics\/orbital-power-grid-space-computing-data-centers\" target=\"_blank\" rel=\"dofollow noopener\">infrastructure <\/a>by the end of 2026. The company said the chip is the first generation of a multigenerational roadmap, with second- and third-generation designs already in development.<\/p>\n<p class=\"wp-block-paragraph\">The broader goal is to make inference faster and more power-efficient as <a href=\"https:\/\/interestingengineering.com\/innovation\/openai-new-data-center-nvidia\" target=\"_blank\" rel=\"dofollow noopener\">OpenAI<\/a> expands its computing infrastructure and serves increasingly capable models.<\/p>\n","protected":false},"excerpt":{"rendered":"OpenAI says its first custom inference chip, Jalape\u00f1o, can deliver more AI work per watt while also reducing&hellip;\n","protected":false},"author":2,"featured_media":151373,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[4518,73873,73874,3887,55677,59615,157,73858,73875],"class_list":["post-151372","post","type-post","status-publish","format-standard","has-post-thumbnail","category-openai","tag-ai-chips","tag-custom-inference-chip","tag-gpt-oss","tag-inference","tag-kimi-k2-5","tag-latency","tag-openai","tag-openai-jalapeno","tag-performance-per-watt"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/151372","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=151372"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/151372\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/151373"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=151372"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=151372"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=151372"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}