{"id":80989,"date":"2026-06-21T17:19:14","date_gmt":"2026-06-21T17:19:14","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/80989\/"},"modified":"2026-06-21T17:19:14","modified_gmt":"2026-06-21T17:19:14","slug":"building-a-dense-agentic-ai-cpu-rack-today-page-2-of-3","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/80989\/","title":{"rendered":"Building a Dense Agentic AI CPU Rack Today &#8211; Page 2 of 3"},"content":{"rendered":"<p>            What Makes Good CPUs for Agentic AI?<\/p>\n<p>CPU performance starts with instructions per clock, clock speed, and core count. IPC determines how much work a core can do per cycle. That depends on core architecture, on-chip cache, and memory access speed. If the cores cannot get data fast enough, the whole system waits on memory. Fast memory access helps with initial loads and any fetches. PCIe to NICs and storage sit roughly at on-chip speeds on modern CPUs, though individual implementations vary.<\/p>\n<p><a href=\"https:\/\/www.servethehome.com\/building-a-dense-agentic-ai-cpu-rack-amd-dell-today\/agentsth-lmbench-4k-amd-epyc-8125p-and-8324p\/\" rel=\"attachment wp-att-101312 nofollow noopener\" target=\"_blank\"><img fetchpriority=\"high\" decoding=\"async\" class=\"size-large wp-image-101312\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/AgentSTH-LMbench-4K-AMD-EPYC-8125P-and-8324P--800x398.jpeg\" alt=\"AgentSTH LMbench 4K AMD EPYC 8125P And 8324P\" width=\"696\" height=\"346\"  \/><\/a>AgentSTH LMbench 4K AMD EPYC 8125P And 8324P<\/p>\n<p>Higher clock speeds help in two ways. First, the time to the next clock refresh, where a task can start, is shorter. Second, if a job is a fixed number of instructions, more cycles per second means faster answers. More cores mean more cores cycling, which means more requests can be serviced simultaneously.<\/p>\n<p>At a high level, CPUs take data and instructions and turn them over millions of times per second per core. The goal for running AI agents is to get them onto the fastest cores available and scale out across many agents.<\/p>\n<p><a href=\"https:\/\/www.servethehome.com\/building-a-dense-agentic-ai-cpu-rack-amd-dell-today\/agentsth-core-to-core-latency-amd-epyc-8125p-and-8324p\/\" rel=\"attachment wp-att-101311 nofollow noopener\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"size-large wp-image-101311\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/AgentSTH-Core-to-Core-Latency-AMD-EPYC-8125P-and-8324P--800x317.jpeg\" alt=\"AgentSTH Core To Core Latency AMD EPYC 8125P And 8324P\" width=\"696\" height=\"276\"  \/><\/a>AgentSTH Core To Core Latency AMD EPYC 8125P And 8324P<\/p>\n<p>Once you start running real agentic AI workflows, you quickly see that very few workloads actually use all cores at once. Or better said, on today\u2019s modern multi-core CPUs, the performance of an individual agentic AI workload is not really running workload X over 64, 128, or 192 cores per system. Instead, the likely case is that both the side issuing commands (the AI agent side) and the side receiving commands (the application side) are running many CPU cores on different workloads simultaneously. With AgentSTH, not only are we looking at the entire socket and single-core performance, but we are also running Agentic workloads spanning different numbers of cores. To be clear, there are actually workloads that, if you use them across an entire CPU, run slower because they are completely single-thread limited. Having a massive number of cores waiting for one core to finish is silly.<\/p>\n<p><a href=\"https:\/\/www.servethehome.com\/amd-epyc-genoa-gaps-intel-xeon-in-stunning-fashion\/amd-epyc-9654-2p-384-thread-htop-1-core\/\" rel=\"attachment wp-att-65403 nofollow noopener\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"size-large wp-image-65403\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/AMD-EPYC-9654-2P-384-Thread-HTOP-1-core-800x192.jpg\" alt=\"AMD EPYC 9654 2P 384 Thread HTOP 1 Core\" width=\"696\" height=\"167\"  \/><\/a>AMD EPYC 9654 2P 384 Thread HTOP 1 Core<\/p>\n<p>One way folks have traditionally looked at this is by looking at things like single-core performance.<\/p>\n<p><a href=\"https:\/\/www.servethehome.com\/building-a-dense-agentic-ai-cpu-rack-amd-dell-today\/agentsth-v7-single-core-dimensions-amd-epyc-8125p-and-8324p\/\" rel=\"attachment wp-att-101313 nofollow noopener\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"size-large wp-image-101313\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/AgentSTH-V7-Single-Core-Dimensions-AMD-EPYC-8125P-and-8324P--800x282.jpeg\" alt=\"AgentSTH V7 Single Core Dimensions AMD EPYC 8125P And 8324P\" width=\"696\" height=\"245\"  \/><\/a>AgentSTH V7 Single Core Dimensions AMD EPYC 8125P And 8324P<\/p>\n<p>That has led to a lot of interesting philosophical questions in the industry around how to measure performance. With AgentSTH V7 we are using throughput, coordination, and memory bandwidth as our high-level dimensions of CPU performance, but there are a few worth watching out for.<\/p>\n<p>Performance Per Core and Memory Bandwidth<\/p>\n<p>Performance per core is a strangely defined metric. AMD and Intel both define a core as two threads on one physical core, and that distinction gets abused in marketing. Single-core performance often means a benchmark running one copy on one core, which does not leverage the SMT thread and can leave 30 percent or more of the performance on the table. Performance per thread divides total performance by total thread count, which distorts the picture because an SMT thread is typically a 20 to 30 percent adder, not a full core. Sockets dominate infrastructure cost because each one goes into a machine with NICs, SSDs, chassis, motherboard, and power delivery.<\/p>\n<p><a href=\"https:\/\/www.servethehome.com\/building-a-dense-agentic-ai-cpu-rack-amd-dell-today\/sth-what-counts-as-a-core\/\" rel=\"attachment wp-att-101323 nofollow noopener\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"size-large wp-image-101323\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/STH-What-Counts-as-a-Core-800x447.jpeg\" alt=\"STH What Counts As A Core\" width=\"696\" height=\"389\"  \/><\/a>STH What Counts As A Core<\/p>\n<p>Per-core performance is very important, without a doubt. That is what gates many agentic workflows as well as legacy applications. At the same time, performance per socket is also hugely important. Sockets affect the number of NICs, motherboards, SSDs, and so forth required for a system. The more we looked into it, the more we kept coming back to simple advice. You want as many fast CPU cores per socket and per rack as you can get for the agentic AI era. Lots of slow cores are probably not as useful because they might be too slow. Few fast cores are interesting, but you end up needing to pack a lot of sockets into a rack to scale that model. Having fewer cores per socket has one advantage you hear a lot about. The same number of memory channels and fewer cores means you have more memory bandwidth per core, and that is often why many HPC workloads actually favor lower core count parts.<\/p>\n<p>Memory bandwidth is a major factor. We now have an entire bucket for mostly memory-bandwidth-sensitive workloads in AgentSTH V7. Some agentic AI workloads, such as in-memory databases and HPC workloads, are heavily memory-bandwidth-bound. We found this in several AgentSTH benchmark sub-tests during months of profiling. STREAM remains the industry\u2019s go-to memory bandwidth benchmark, and it shows up in marketing for every new processor generation as the generational high-water mark. Companies commonly use the STREAM Triad score, though some include additional tests in a geometric mean to further boost the new generation. Note, I am not talking about one company doing this. Everyone does it because it is a high-water mark.<\/p>\n<p><a href=\"https:\/\/www.servethehome.com\/building-a-dense-agentic-ai-cpu-rack-amd-dell-today\/amd-epyc-milan-to-genoa-example-ddr4-to-ddr5\/\" rel=\"attachment wp-att-101324 nofollow noopener\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"size-large wp-image-101324\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/AMD-EPYC-Milan-to-Genoa-Example-DDR4-to-DDR5-800x451.jpeg\" alt=\"AMD EPYC Milan To Genoa Example DDR4 To DDR5\" width=\"696\" height=\"392\"  \/><\/a>AMD EPYC Milan To Genoa Example DDR4 To DDR5<\/p>\n<p>Memory bandwidth roughly equals the speed of each memory module multiplied by the number of channels. Moving from AMD EPYC 7003 Milan with 8-channel DDR4-3200 to EPYC 9004 Genoa with 12-channel DDR5-4800 gives you 57600 divided by 25600, which is 2.25 times the memory bandwidth per socket. That makes STREAM scores go up significantly, even if IPC does not improve as much per core. It also means the geometric means of benchmark results trend higher when one or more tests are memory bandwidth bound.\u00a0That is generally. If you want to see this in action,\u00a0<a href=\"https:\/\/www.servethehome.com\/here-is-why-you-should-fully-populate-memory-channels-on-cpus-featuring-amd-epyc-genoa\/\" target=\"_blank\" rel=\"noopener nofollow\">here is Why You Should Fully Populate Memory Channels on CPUs featuring AMD EPYC Genoa<\/a>.\u00a0We\u00a0showed the scaling as we scaled the number of installed DIMMs:<\/p>\n<p><a href=\"https:\/\/www.servethehome.com\/here-is-why-you-should-fully-populate-memory-channels-on-cpus-featuring-amd-epyc-genoa\/screenshot-24\/\" rel=\"attachment wp-att-80916 nofollow noopener\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"size-large wp-image-80916\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/AMD-EPYC-Stream-Triad-by-Number-of-Memory-Channels-Filled-Total-Bandwidth-800x511.jpg\" alt=\"AMD EPYC Stream Triad by Number of Memory Channels Filled Bandwidth\" width=\"696\" height=\"445\"  \/><\/a>AMD EPYC Stream Triad by Number of Memory Channels Filled Bandwidth<\/p>\n<p>If you want to see the pattern, here is what the incremental bandwidth per DIMM we got when adding more memory. That is a fairly flat line showing why STREAM largely scales with the number of memory channels and the throughput of each device.<\/p>\n<p><a href=\"https:\/\/www.servethehome.com\/here-is-why-you-should-fully-populate-memory-channels-on-cpus-featuring-amd-epyc-genoa\/screenshot-26\/\" rel=\"attachment wp-att-80918 nofollow noopener\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"size-large wp-image-80918\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/AMD-EPYC-Stream-Triad-by-Number-of-Memory-Channels-Filled-Average-Bandwidth-Per-DIMM-1-800x508.jpg\" alt=\"AMD EPYC Stream Triad by Number of Memory Channels Filled Average Bandwidth Per DIMM\" width=\"696\" height=\"442\"  \/><\/a>AMD EPYC Stream Triad by Number of Memory Channels Filled Average Bandwidth Per DIMM<\/p>\n<p>Next-generation processors will support DDR5-8000 to DDR5-12800 with up to 16 memory channels. Most current systems have 12 channels at DDR5-6400. Moving to 16 channels at DDR5-8000 gives 128000 divided by 76800, which is 1.667 times faster on certain workloads. With 16 channels at DDR5-12800, which reaches 204800 divided by 76800, or 2.667 times faster. Next-gen CPUs using MRDIMMs instead of LPDDR5X can deliver substantial gains in memory bandwidth on the right workloads. That is going to be across vendors, but it is mostly just a function of having faster memory and more channels in the new generation.<\/p>\n<p><a href=\"https:\/\/www.servethehome.com\/building-a-dense-agentic-ai-cpu-rack-amd-dell-today\/late-2026-2027-era-mainstream-server-cpu-memory-bandwidth-increases\/\" rel=\"attachment wp-att-101329 nofollow noopener\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"size-large wp-image-101329\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/Late-2026-2027-Era-Mainstream-Server-CPU-Memory-Bandwidth-Increases-800x447.jpeg\" alt=\"Late 2026 2027 Era Mainstream Server CPU Memory Bandwidth Increases\" width=\"696\" height=\"389\"  \/><\/a>Late 2026 2027 Era Mainstream Server CPU Memory Bandwidth Increases<\/p>\n<p>Agentic AI workflows have legitimate areas that will be memory bandwidth-limited as tools get called. Compression workloads can be memory bandwidth-bound as well, and that is one thing folks often overlook. Other areas are CPU-bound. Compiling source code is often CPU-bound. Legacy applications with databases face CPU constraints and potentially per-core licensing costs. The interesting part is that, on the agentic AI CPU side, there is a net-new workload that looks like the traditional server CPU compute side, albeit with slight adjustments to the weighting.<\/p>\n<p><a href=\"https:\/\/www.servethehome.com\/amd-epyc-9005-turin-turns-transcendent-performance-solidigm-broadcom\/screenshot-91\/\" rel=\"attachment wp-att-81526 nofollow noopener\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"size-large wp-image-81526\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/AMD-EPYC-9965-STREAM-192C-800x686.jpg\" alt=\"AMD EPYC 9965 STREAM 192C\" width=\"696\" height=\"597\"  \/><\/a>AMD EPYC 9965 STREAM 192C<\/p>\n<p>For STH readers, a strong consulting engagement would be moving companies to newer, faster servers or off per-core license databases, since agentic AI is going to put more pressure on existing applications. Next however, let us quickly take a look at what \u201cgood\u201d looks like if you are deploying today.<\/p>\n","protected":false},"excerpt":{"rendered":"What Makes Good CPUs for Agentic AI? CPU performance starts with instructions per clock, clock speed, and core&hellip;\n","protected":false},"author":2,"featured_media":79926,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[179,7493,43470,3125,18486,43471,43472],"class_list":["post-80989","post","type-post","status-publish","format-standard","has-post-thumbnail","category-agentic-ai","tag-agentic-ai","tag-agentic-artificial-intelligence","tag-agentsth","tag-amd","tag-dell","tag-epyc","tag-turin"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/80989","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=80989"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/80989\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/79926"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=80989"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=80989"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=80989"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}