{"id":107700,"date":"2026-07-16T04:17:07","date_gmt":"2026-07-16T04:17:07","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/107700\/"},"modified":"2026-07-16T04:17:07","modified_gmt":"2026-07-16T04:17:07","slug":"agentic-inference-reshapes-ai-infrastructure-at-raise-summit","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/107700\/","title":{"rendered":"Agentic inference reshapes AI infrastructure at RAISE Summit"},"content":{"rendered":"<p>Agentic inference is reshaping the center of gravity in <a href=\"https:\/\/siliconangle.com\/2026\/06\/26\/raise-summit-ai-infrastructure-thecube-raisesummit\/\" rel=\"nofollow noopener\" target=\"_blank\">AI infrastructure<\/a>. What began as a race to scale training has shifted into a phase defined by expanding context windows, memory\u2011augmented reasoning and the need to keep graphics processing units continuously fed with data.<\/p>\n<p>As enterprises push deeper into agentic systems, storage has moved into the critical path of AI performance. Agentic inference is exposing new bottlenecks, accelerating the adoption of design patterns and pushing organizations to rethink how intelligence is staged, retrieved and delivered to GPUs, according to <a href=\"https:\/\/www.linkedin.com\/in\/greg-matson-102b12\/\" rel=\"nofollow noopener\" target=\"_blank\">Greg Matson<\/a>\u00a0(pictured), senior vice president and head of marketing and products at Solidigm, a trademark of SK Hynix NAND Product Solutions Corp.<\/p>\n<p>\u201cIt started a couple of years ago with training, where the need for high-capacity, high-performance storage very adjacent to the GPUs was all of a sudden center stage,\u201d Matson told theCUBE. \u201cBut now, as we go from last year to this year, inference phase into agentic inference, it\u2019s exploding even more. Storage is actually a whole new storage tier that\u2019s being created to extend the memory for the system.\u201d<\/p>\n<p>Matson spoke with theCUBE at <a href=\"https:\/\/siliconangle.com\/tag\/raisesummit26eventpage\/\" rel=\"nofollow noopener\" target=\"_blank\">RAISE Summit<\/a>, during an exclusive broadcast on theCUBE, SiliconANGLE Media\u2019s livestreaming studio. Conversations with cross-industry leaders showed how agentic inference is expanding infrastructure design beyond raw compute to encompass specialized architectures, memory and storage, capital deployment and sovereign control of enterprise data. (* Disclosure below.)<\/p>\n<p>Here\u2019s theCUBE\u2019s complete video interview with Greg Matson:<\/p>\n<p>Here are three insights you may have missed from theCUBE\u2019s coverage of <a href=\"https:\/\/siliconangle.com\/tag\/raisesummit26eventpage\/\" rel=\"nofollow noopener\" target=\"_blank\">RAISE Summit<\/a>:<\/p>\n<p>Insight #1: Agentic inference drives specialization across the AI stack.<\/p>\n<p>Advanced Micro Devices Inc. is responding to varied AI workloads by optimizing across CPUs, GPUs, adaptive computing and networking rather than focusing on individual chips. Its <a href=\"https:\/\/www.amd.com\/en\/products\/software\/rocm.html\" rel=\"nofollow noopener\" target=\"_blank\">ROCm software<\/a> stack is designed to provide a consistent layer across data center clusters, edge deployments and AI-enabled PCs, according to <a href=\"https:\/\/www.linkedin.com\/in\/mark-papermaster-66914925\/\" rel=\"nofollow noopener\" target=\"_blank\">Mark Papermaster<\/a>, chief technology officer and executive vice president of AMD.<\/p>\n<p>\u201cThe workloads are so complex because people are looking at what they do end to end,\u201d he said during the event. \u201cThey\u2019re looking at whole processes, not just one bespoke task. That means you need different computing engines, and they need to <a href=\"https:\/\/siliconangle.com\/2026\/07\/08\/ai-infrastructure-optimization-raisesummit\/\" rel=\"nofollow noopener\" target=\"_blank\">work together at scale<\/a>. We\u2019re talking across massive clusters of racks.\u201d<\/p>\n<p>Tensordyne Inc. is addressing power constraints by changing the math inside the silicon. Its <a href=\"https:\/\/www.tensordyne.ai\/\" rel=\"nofollow noopener\" target=\"_blank\">Napier inference chip<\/a> uses the proprietary Pareto logarithmic number system, replacing multiplications with additions to reduce reliance on large, power-intensive multiplier circuits. A 72-chip Napier pod draws 30 kilowatts, compared with 150 kilowatts for a comparable Nvidia Corp. system, according to <a href=\"https:\/\/www.linkedin.com\/in\/gillesbackhus\/\" rel=\"nofollow noopener\" target=\"_blank\">Gilles Backhus<\/a>, co-founder of Tensordyne.<\/p>\n<p>\u201cOur logarithmic math \u2014 it\u2019s completely under the hood,\u201d he told theCUBE. \u201cFrom a user point of view, from a [software development kit] point of view, you don\u2019t even notice it. It just looks like normal floating-point math. It\u2019s just that the engine under the hood <a href=\"https:\/\/siliconangle.com\/2026\/07\/08\/logarithmic-math-accelerates-enterprise-ai-inference-raisesummit\/\" rel=\"nofollow noopener\" target=\"_blank\">is more efficient<\/a>.\u201d<\/p>\n<p>Purpose-built accelerators are also moving into production alongside GPUs. A <a href=\"https:\/\/www.prnewswire.com\/news-releases\/parasail-to-combine-nvidia-ai-infrastructure-with-d-matrix-accelerators-to-achieve-10x-faster-token-generation-302820178.html\" rel=\"nofollow noopener\" target=\"_blank\">Parasail deployment<\/a> pairs d-Matrix <a href=\"https:\/\/www.d-matrix.ai\/product\/\" rel=\"nofollow noopener\" target=\"_blank\">Corsair accelerators<\/a> with Nvidia Hopper and Blackwell GPUs to serve the different requirements of compute-heavy prefill and latency-sensitive token generation, according to d-Matrix Corp. co-founders <a href=\"https:\/\/www.linkedin.com\/in\/sudeep-bhoja-070a111\/\" rel=\"nofollow noopener\" target=\"_blank\">Sudeep Bhoja<\/a>, chief technology officer, and <a href=\"http:\/\/linkedin.com\/in\/sheth\" rel=\"nofollow noopener\" target=\"_blank\">Sid Sheth<\/a>, president and chief executive officer. It represents an early commercial-scale example of heterogeneous inference in production.<\/p>\n<p>\u201cLow latency is the name of the game today,\u201d Bhoja said during the event. \u201cAgents are running for a long time; users don\u2019t want to wait. So, there\u2019s a demand on trying to get the latency down, and that means <a href=\"https:\/\/siliconangle.com\/2026\/07\/09\/fast-token-generation-accelerates-enterprise-ai-inference-raisesummit\/\" rel=\"nofollow noopener\" target=\"_blank\">disaggregated inference<\/a>.\u201d<\/p>\n<p>Here\u2019s the complete video interview with Sudeep Bhoja and Sid Sheth:<\/p>\n<p>Insight #2: Storage becomes an active extension of AI memory.<\/p>\n<p>As agentic inference expands from individual prompts to longer-running sessions, the <a href=\"https:\/\/siliconangle.com\/2026\/07\/07\/storage-technology-value-agentic-ai-raisesummit\/\" rel=\"nofollow noopener\" target=\"_blank\">volume of context data<\/a> can exceed GPU memory capacity. Hyperscalers are replacing legacy infrastructure with high-capacity solid-state storage positioned near accelerators, according to Matson.<\/p>\n<p>\u201cWhile the GPU is the most expensive part of your infrastructure, you want that thing to be humming 100% of the time generating tokens,\u201d he told theCUBE. \u201cAnd if it\u2019s down, waiting for data, then you\u2019re <a href=\"https:\/\/siliconangle.com\/2026\/07\/08\/intelligence-layer-memory-extension-accelerates-agentic-ai-raisesummit\/\" rel=\"nofollow noopener\" target=\"_blank\">wasting your money on GPUs<\/a>.\u201d<\/p>\n<p>Solidigm is testing storage as part of complete AI systems rather than as an isolated component. Its <a href=\"https:\/\/news.solidigm.com\/en-WW\/254753-solidigm-unveils-ai-central-lab-home-to-highest-performing-and-most-dense-storage-test-clusters\/\" rel=\"nofollow noopener\" target=\"_blank\">AI Central Lab<\/a> runs actual workloads across accelerator hardware and partner software. At the same time, high-density solid-state drive configurations show how greater capacity in less rack space can reduce storage power demands, according to <a href=\"https:\/\/www.linkedin.com\/in\/avishetty\/\" rel=\"nofollow noopener\" target=\"_blank\">Avi Shetty<\/a>, vice president of AI ecosystem, solutions and market enablement at Solidigm.<\/p>\n<p>\u201cNobody cares about random read, random write on a data sheet right now,\u201d Shetty said. \u201cWhat people care about is how does it operate in an <a href=\"https:\/\/siliconangle.com\/2026\/07\/09\/token-per-watt-metrics-optimize-ai-data-center-efficiency-raisesummit\/\" rel=\"nofollow noopener\" target=\"_blank\">AI data center<\/a> and an AI workload.\u201d<\/p>\n<p>Here\u2019s theCUBE\u2019s complete video interview with Avi Shetty:<\/p>\n<p>Insight #3: Capital and sovereignty become part of the AI infrastructure stack.<\/p>\n<p>Agentic inference projects require more than power and GPUs; they can stall when developers can\u2019t assemble financing quickly enough. Argentum AI Inc.\u2019s demand-first model secures customers <a href=\"https:\/\/siliconangle.com\/2026\/07\/08\/capital-stack-financing-accelerates-ai-deployment-raisesummit\/\" rel=\"nofollow noopener\" target=\"_blank\">before committing capital<\/a>, using contracted revenue to support construction while remaining neutral regarding silicon and original equipment manufacturers, according to <a href=\"https:\/\/www.linkedin.com\/in\/andrew-sobko-662231102\/\" rel=\"nofollow noopener\" target=\"_blank\">Andrew Sobko<\/a>, founder and chief executive officer of Argentum AI Inc.<\/p>\n<p>\u201cWe formed the view that the biggest bottleneck in the focus on the speed of deployment is a capital stack,\u201d he told theCUBE. \u201cHow do you get the projects financed as fast as possible? That sort of became one of our core products, where we call it bringing power, compute and capital.\u201d<\/p>\n<p>Data sovereignty, meanwhile, is moving from compliance into <a href=\"https:\/\/siliconangle.com\/2026\/07\/09\/data-sovereignty-models-protect-proprietary-enterprise-ai-raisesummit\/\" rel=\"nofollow noopener\" target=\"_blank\">infrastructure architecture<\/a> as agentic inference draws on proprietary enterprise context. The concept spans territorial, operational, stack, legal and unit economics concerns, according to <a href=\"https:\/\/www.linkedin.com\/in\/amit-eyal-govrin-4a045b312\/\" rel=\"nofollow noopener\" target=\"_blank\">Amit Eyal Govrin<\/a>, chief executive officer of Agentcy Labs Inc., and <a href=\"http:\/\/linkedin.com\/in\/prathle\" rel=\"nofollow noopener\" target=\"_blank\">Philip Rathle<\/a>, chief technology officer of Neo4j Inc.<\/p>\n<p>\u201cSovereignty is exerting agency and control over your AI,\u201d Govrin said during the event. \u201cYou have to be free and clear of state, economic and threat actors overtaking any level of control over your stack. You\u2019re not paying rent to somebody else.\u201d<\/p>\n<p>Knowledge graphs add another form of control by allowing some decisions to run deterministically rather than relying entirely on probabilistic models. That can give enterprises consistent business rules, explainability and governance alongside the flexibility of large language models, Rathle noted.<\/p>\n<p>\u201cHaving the capacity to do AI with a full brain, both hemispheres, is ultra important,\u201d he told theCUBE. \u201cLLMs are spontaneous, creative \u2014 they make mistakes, you don\u2019t know why. Having the graph as the left brain to the LLM right brain is really at the core of where graphs fit in.\u201d<\/p>\n<p>Here\u2019s theCUBE\u2019s complete video interview with Amit Eyal Govrin and Philip Rathle:<\/p>\n<p>To watch more of theCUBE\u2019s coverage of RAISE Summit, here\u2019s our <a href=\"https:\/\/www.youtube.com\/playlist?list=PLYzJZhQ_K43U\" rel=\"nofollow noopener\" target=\"_blank\">complete video playlist<\/a>:<\/p>\n<p>(* Disclosure: TheCUBE is a paid media partner for the RAISE Summit event. Neither Solidigm, the headline sponsor of theCUBE\u2019s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)<\/p>\n<p>Photo: SiliconANGLE<\/p>\n<p>Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE\u2019s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.<\/p>\n<p>15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more<br \/>\n11.4k+ theCUBE alumni \u2014 Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.<\/p>\n<p>About SiliconANGLE Media<\/p>\n<p>SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of <a href=\"https:\/\/cts.businesswire.com\/ct\/CT?id=smartlink&amp;url=https%3A%2F%2Fsiliconangle.com%2F&amp;esheet=54119777&amp;newsitemid=20240910506833&amp;lan=en-US&amp;anchor=SiliconANGLE&amp;index=9&amp;md5=646b1b564e2259100a2b8638aab0a552\" rel=\"nofollow noopener\" target=\"_blank\">SiliconANGLE<\/a>, <a href=\"https:\/\/cts.businesswire.com\/ct\/CT?id=smartlink&amp;url=https%3A%2F%2Fwww.thecube.net%2F&amp;esheet=54119777&amp;newsitemid=20240910506833&amp;lan=en-US&amp;anchor=theCUBE+Network&amp;index=10&amp;md5=7de2a85f95ab4a4a495cede20b8cb1da\" rel=\"nofollow noopener\" target=\"_blank\">theCUBE Network<\/a>, <a href=\"https:\/\/cts.businesswire.com\/ct\/CT?id=smartlink&amp;url=https%3A%2F%2Fthecuberesearch.com%2F&amp;esheet=54119777&amp;newsitemid=20240910506833&amp;lan=en-US&amp;anchor=theCUBE+Research&amp;index=11&amp;md5=7bb33676722925eb57d588ec343e4f6f\" rel=\"nofollow noopener\" target=\"_blank\">theCUBE Research<\/a>, <a href=\"https:\/\/cts.businesswire.com\/ct\/CT?id=smartlink&amp;url=https%3A%2F%2Fwww.cube365.net%2F&amp;esheet=54119777&amp;newsitemid=20240910506833&amp;lan=en-US&amp;anchor=CUBE365&amp;index=12&amp;md5=d310fb35919714e66ad8d42c9c0c1bc6\" rel=\"nofollow noopener\" target=\"_blank\">CUBE365<\/a>, <a href=\"https:\/\/cts.businesswire.com\/ct\/CT?id=smartlink&amp;url=https%3A%2F%2Fwww.thecubeai.com%2F&amp;esheet=54119777&amp;newsitemid=20240910506833&amp;lan=en-US&amp;anchor=theCUBE+AI&amp;index=13&amp;md5=b8b98472f8071b23ebb10ab9a8dd0683\" rel=\"nofollow noopener\" target=\"_blank\">theCUBE AI<\/a> and theCUBE SuperStudios \u2014 with flagship locations in Silicon Valley and the New York Stock Exchange \u2014 SiliconANGLE Media operates at the intersection of media, technology and AI.<\/p>\n<p>Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.<\/p>\n","protected":false},"excerpt":{"rendered":"Agentic inference is reshaping the center of gravity in AI infrastructure. What began as a race to scale&hellip;\n","protected":false},"author":2,"featured_media":107701,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[179,7493,396,55710,45672],"class_list":["post-107700","post","type-post","status-publish","format-standard","has-post-thumbnail","category-agentic-ai","tag-agentic-ai","tag-agentic-artificial-intelligence","tag-siliconangle","tag-three-insights-you-may-have-missed-from-thecubes-coverage-of-raise-summit","tag-victoria-gayton"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/107700","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=107700"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/107700\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/107701"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=107700"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=107700"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=107700"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}