{"id":151039,"date":"2026-08-25T19:43:07","date_gmt":"2026-08-25T19:43:07","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/151039\/"},"modified":"2026-08-25T19:43:07","modified_gmt":"2026-08-25T19:43:07","slug":"ai-storage-infrastructure-supports-scalable-ai-inference-siliconangle-ai-inference-gets-a-new-tier-as-context-windows-grow","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/151039\/","title":{"rendered":"AI storage infrastructure supports scalable AI inference &#8211; SiliconANGLE AI inference gets a new tier as context windows grow"},"content":{"rendered":"<p>AI storage infrastructure is becoming a more consequential planning issue as organizations move from model training toward agentic AI. As agents reason, act and reassess, they build longer contexts and generate more data that they <a href=\"https:\/\/siliconangle.com\/2026\/08\/19\/ai-inference-infrastructure-requires-full-stack-coordination-supermicroopenstoragesummit\/\" rel=\"nofollow noopener\" target=\"_blank\">must access quickly during inference<\/a>.<\/p>\n<p>Agentic AI is also changing the shape of the data problem. Interactions are growing longer and producing more information. At the same time, the data\u2019s size, importance and movement through the infrastructure can all affect how quickly an application responds, according to <a href=\"https:\/\/www.linkedin.com\/in\/scottshadley\/\" rel=\"nofollow noopener\" target=\"_blank\">Scott Shadley<\/a> (pictured, left), director of technology planning at Solidigm Inc.<\/p>\n<p>\u201cOne of the beautiful things that\u2019s happened in this agentic AI, or even just the AI era, is [that] people are starting to pay attention to storage,\u201d he said. \u201cWhat\u2019s unique about this particular era is it\u2019s no longer one- or two-dimensional. We have data magnitude and growth in size, importance and all of the other volumetric aspects of that. But at the end of the day, it comes down to that bit of data and how fast that bit of data moves from point A to point B.\u201d<\/p>\n<p>Shadley, along with <a href=\"https:\/\/www.linkedin.com\/in\/anat-heilper\/\" rel=\"nofollow noopener\" target=\"_blank\">Anat Heilper<\/a> (center), director of AI architecture at Vast Data Inc., and <a href=\"https:\/\/www.linkedin.com\/in\/benlee73\/\" rel=\"nofollow noopener\" target=\"_blank\">Ben Lee<\/a> (right), director of solution management at Super Micro Computer Inc., spoke with theCUBE Research\u2019s <a href=\"https:\/\/www.linkedin.com\/in\/robstrechay\/\" rel=\"nofollow noopener\" target=\"_blank\">Rob Strechay<\/a> during the <a href=\"https:\/\/www.thecube.net\/events\/supermicro\/open-storage-summit-2026?utm_source=siliconangle&amp;utm_medium=article&amp;utm_campaign=supermicro_open_storage_summit_2026\" rel=\"nofollow noopener\" target=\"_blank\">Supermicro Open Storage Summit interview series<\/a>. They discussed how the companies\u2019 respective technologies fit together as expanding context windows and KV caches create new storage and memory demands for agentic AI.<\/p>\n<p>AI storage infrastructure brings KV cache closer to compute<\/p>\n<p>As context windows expand, graphics processing unit memory alone can\u2019t hold everything an agentic workload needs during inference. AI storage infrastructure must therefore provide additional tiers that balance proximity, capacity and speed, with each layer handling a different part of the data load, according to Shadley.<\/p>\n<p>\u201cAs you think through that architecture \u2014 and you need to put that context somewhere \u2014 that context can start living in what used to be a no-no zone,\u201d he said. \u201cSo [solid-state drives] have found a new home. One of the unique things about this 3.5 tier that we\u2019re creating is that it could not exist until we had things like [Non-Volatile Memory Express] SSDs.\u201d<\/p>\n<p>That middle tier is one part of a larger architecture. Solidigm supplies <a href=\"https:\/\/www.solidigm.com\/\" rel=\"nofollow noopener\" target=\"_blank\">SSDs<\/a> that provide fast access to cached data, Supermicro integrates them into <a href=\"https:\/\/www.supermicro.com\/en\/solutions\/rack-integration\" rel=\"nofollow noopener\" target=\"_blank\">rack-scale systems<\/a> and Vast Data\u2019s <a href=\"https:\/\/www.vastdata.com\/platform\/ai-os\" rel=\"nofollow noopener\" target=\"_blank\">AI Operating System<\/a> provides network storage and the data services needed to use, manage and protect that data alongside broader AI workloads, Heilper noted.<\/p>\n<p>\u201cWhen we talk about [key-value] cache, which is a very significant optimization that can be done in AI inferencing \u2026 in essence, it\u2019s the ability to replace compute with storage,\u201d she said. \u201cThis is very significant because we all know the GPU is very, very expensive. When you have very high KV cache hit rates, we both save on compute and reduce the latency significantly.\u201d<\/p>\n<p>AI infrastructure needs room to evolve<\/p>\n<p>Supermicro\u2019s Context Memory eXtension, or CMX, proposal targets organizations with large AI clusters and substantial data demands; other deployments may require different combinations of memory, local SSDs and network storage. The company\u2019s broader value lies in composing those building blocks around each customer\u2019s workload, rather than treating one architecture as a universal answer, according to Lee.<\/p>\n<p>\u201cWe believe solving the problem will take the whole rack because you cannot just buy more GPUs with more [high-bandwidth memory] \u2026 it\u2019s very expensive,\u201d he said. \u201cAll the KV cache will naturally overflow from the GPU, HBM, to the system memory, to the local SSD and to the network storage. But there\u2019s a new industrial definition to try and fill the gap, and they call it G3.5, which is the CMX solution. This is a very AI-native KV cache tier that can fulfill the demand.\u201d<\/p>\n<p>Vast Data has been testing KV cache offload with partners across different software environments. Those experiments aim to show how the architecture behaves when cached context is integrated into production inference workloads, according to Heilper.<\/p>\n<p>\u201cWith Nvidia <a href=\"https:\/\/www.nvidia.com\/en-us\/ai\/dynamo\/\" rel=\"nofollow noopener\" target=\"_blank\">Dynamo<\/a>, we\u2019ve shown that we managed to get 20 times faster time-to-first-token, which means that the latency that you perceive as a user is significantly faster,\u201d she said. \u201cAlso, [we\u2019ve] seen 90% savings in GPU time.\u201d<\/p>\n<p>Those results depend on an AI storage infrastructure that can match storage performance and capacity to the workload. Solidigm\u2019s D7-PS1010 performance-oriented SSD and D5-P5336 capacity drive address different points in the hierarchy, according to Shadley. Supermicro integrates those components into systems that can be configured around customer requirements.<\/p>\n<p>\u201cThis is not the only definition of a hierarchy stack,\u201d Shadley said. \u201cIt\u2019s the current primary that everybody leverages as the gold standard, but it\u2019s continuing to evolve, and it\u2019s unique. Being very proactive with your customer or your supplier to better understand what they know about what you need is no longer transactional. Those value-level transaction conversations are now what are going to drive the future of us deploying these types of architectures.\u201d<\/p>\n<p>Stay tuned for the complete video, part of SiliconANGLE and theCUBE\u2019s coverage of the <a href=\"https:\/\/www.thecube.net\/events\/supermicro\/open-storage-summit-2026?utm_source=siliconangle&amp;utm_medium=article&amp;utm_campaign=supermicro_open_storage_summit_2026\" rel=\"nofollow noopener\" target=\"_blank\">Supermicro Open Storage Summit interview series<\/a>.<\/p>\n<p>(* Disclosure: TheCUBE is a paid media partner for the Supermicro Open Storage Summit interview series. Neither Supermicro, the sponsor of theCUBE\u2019s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)<\/p>\n<p>Photo: SiliconANGLE<\/p>\n<p>Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE\u2019s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.<\/p>\n<p>15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more<br \/>\n11.4k+ theCUBE alumni \u2014 Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network<\/p>\n<p>SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of <a href=\"https:\/\/cts.businesswire.com\/ct\/CT?id=smartlink&amp;url=https%3A%2F%2Fsiliconangle.com%2F&amp;esheet=54119777&amp;newsitemid=20240910506833&amp;lan=en-US&amp;anchor=SiliconANGLE&amp;index=9&amp;md5=646b1b564e2259100a2b8638aab0a552\" rel=\"nofollow noopener\" target=\"_blank\">SiliconANGLE<\/a>, <a href=\"https:\/\/cts.businesswire.com\/ct\/CT?id=smartlink&amp;url=https%3A%2F%2Fwww.thecube.net%2F&amp;esheet=54119777&amp;newsitemid=20240910506833&amp;lan=en-US&amp;anchor=theCUBE+Network&amp;index=10&amp;md5=7de2a85f95ab4a4a495cede20b8cb1da\" rel=\"nofollow noopener\" target=\"_blank\">theCUBE Network<\/a>, <a href=\"https:\/\/cts.businesswire.com\/ct\/CT?id=smartlink&amp;url=https%3A%2F%2Fthecuberesearch.com%2F&amp;esheet=54119777&amp;newsitemid=20240910506833&amp;lan=en-US&amp;anchor=theCUBE+Research&amp;index=11&amp;md5=7bb33676722925eb57d588ec343e4f6f\" rel=\"nofollow noopener\" target=\"_blank\">theCUBE Research<\/a>, <a href=\"https:\/\/cts.businesswire.com\/ct\/CT?id=smartlink&amp;url=https%3A%2F%2Fwww.cube365.net%2F&amp;esheet=54119777&amp;newsitemid=20240910506833&amp;lan=en-US&amp;anchor=CUBE365&amp;index=12&amp;md5=d310fb35919714e66ad8d42c9c0c1bc6\" rel=\"nofollow noopener\" target=\"_blank\">CUBE365<\/a>, <a href=\"https:\/\/cts.businesswire.com\/ct\/CT?id=smartlink&amp;url=https%3A%2F%2Fwww.thecubeai.com%2F&amp;esheet=54119777&amp;newsitemid=20240910506833&amp;lan=en-US&amp;anchor=theCUBE+AI&amp;index=13&amp;md5=b8b98472f8071b23ebb10ab9a8dd0683\" rel=\"nofollow noopener\" target=\"_blank\">theCUBE AI<\/a> and theCUBE SuperStudios \u2014 with flagship locations in Silicon Valley and the New York Stock Exchange \u2014 SiliconANGLE Media operates at the intersection of media, technology and AI.<\/p>\n<p style=\"text-align: left;\">Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.<\/p>\n","protected":false},"excerpt":{"rendered":"AI storage infrastructure is becoming a more consequential planning issue as organizations move from model training toward agentic&hellip;\n","protected":false},"author":2,"featured_media":151040,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[70908,70314,179,24,8158,73741,73742,25,73743,73744,12020,10885,523,10884,655,70912,21685,58,73745,73746,73747,73748,73749,73750,73751,73752,70918,73753,17231,73754,70919,70920,62405,73755],"class_list":["post-151039","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-supermicroopenstoragesummit","tag-thecube","tag-agentic-ai","tag-ai","tag-ai-inference","tag-ai-storage-infrastructure","tag-anat-heilper","tag-artificial-intelligence","tag-ben-lee","tag-cmx","tag-data-infrastructure","tag-data-management","tag-enterprise-ai","tag-enterprise-storage","tag-hbm","tag-hdds","tag-kv-cache","tag-nvidia","tag-nvidia-dynamo","tag-nvme","tag-rack-scale-systems","tag-rob-strechay","tag-scott-shadley","tag-solidigm","tag-solidigm-d7-ps1010","tag-solidigm-ds-p5336","tag-ssds","tag-storage-infrastructure","tag-super-micro-computer","tag-supermicro-g3-5","tag-supermicro-open-storage-summit-2026","tag-supermicroopenstoragesummit26eventpage","tag-vast-data","tag-vast-data-ai-operating-system"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/151039","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=151039"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/151039\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/151040"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=151039"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=151039"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=151039"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}