{"id":333321,"date":"2026-02-12T09:42:16","date_gmt":"2026-02-12T09:42:16","guid":{"rendered":"https:\/\/www.europesays.com\/ie\/333321\/"},"modified":"2026-02-12T09:42:16","modified_gmt":"2026-02-12T09:42:16","slug":"scaling-llama-cpp-on-neoverse-n2-solving-cross-numa-performance-issues","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ie\/333321\/","title":{"rendered":"Scaling llama.cpp On Neoverse N2: Solving Cross-NUMA Performance Issues"},"content":{"rendered":"<p>This blog post explains the cross-NUMA memory access issue that occurs when you run llama.cpp in Neoverse. It also introduces a proof-of-concept patch that addresses this issue and can provide up to a 55% performance increase for text generation when you run the llama3_Q4_0 model on the ZhuFeng Neoverse system.<\/p>\n<p>Cross-NUMA memory access problem<\/p>\n<p>In llama.cpp, performance drops when the number of threads exceeds the number of cores in a NUMA node. This example uses a 64 cores per NUMA node, and the llama3-Q4_0 model.<\/p>\n<p><img data-recalc-dims=\"1\" fetchpriority=\"high\" decoding=\"async\" class=\"alignnone size-full wp-image-24272627\" src=\"https:\/\/www.europesays.com\/ie\/wp-content\/uploads\/2026\/02\/Arm_Scaling-llamacpp-On-Neoverse-N2-fig1.webp.png\" alt=\"\" width=\"1290\" height=\"458\"  \/><\/p>\n<p>Root causes<\/p>\n<p>There are two causes for the problem if threads spawn across NUMA nodes:<\/p>\n<ul>\n<li>The atomic operations in ggml_barrier() across NUMA node is time-consuming.<\/li>\n<li>Lots of cross NUMA memory access of the tensor buffer of Mul Mat operation.<\/li>\n<\/ul>\n<p>ggml_barrier issue<\/p>\n<p>When running llama.cpp in multi-threads, each thread computes part of the tensor data. A barrier is used after all threads finish computation to make sure the data is synced. Performance drops as the number of threads increases and the impact is worse when thread count exceeds a NUMA node (64 cores per node).<\/p>\n<p><img loading=\"lazy\" data-recalc-dims=\"1\" decoding=\"async\" class=\"alignnone size-full wp-image-24272628\" src=\"https:\/\/www.europesays.com\/ie\/wp-content\/uploads\/2026\/02\/Arm_Scaling-llamacpp-On-Neoverse-N2-fig2.webp.png\" alt=\"\" width=\"1294\" height=\"398\"  \/><\/p>\n<p>Mul Mat operation issue<\/p>\n<p>The Mul Mat operation in llama.cpp is the main performance hotspot. In theory, performance should improve as you add more threads. However, this is not the case because the tensor buffer is allocated from malloc() which is not NUMA-aware, and leads to significant cross-NUMA memory access.<\/p>\n<p><img loading=\"lazy\" data-recalc-dims=\"1\" decoding=\"async\" class=\"alignnone size-full wp-image-24272629\" src=\"https:\/\/www.europesays.com\/ie\/wp-content\/uploads\/2026\/02\/Arm_Scaling-llamacpp-On-Neoverse-N2-fig3.webp.png\" alt=\"\" width=\"1324\" height=\"440\"  \/><\/p>\n<p>Optimization methods<\/p>\n<p>To mitigate the cross NUMA problem in this case, two optimization methods are applied:<\/p>\n<ol>\n<li>ggml_barrier optimization<\/li>\n<li>Mul Mat operation optimization<\/li>\n<\/ol>\n<p>ggml_barrier optimization<\/p>\n<p>This solution optimizes the ggml barrier by using the \u201cdivide-and-conquer\u201d approach:<\/p>\n<ul>\n<li>It uses the NUMA local atomic variable to do the barrier which is fast.<\/li>\n<li>Only the last thread of a specified NUMA node syncs with the last threads of other NUMA nodes in a cross-NUMA way.<\/li>\n<\/ul>\n<p>With this method, the number of threads performing cross-NUMA global atomic operations is reduced to the number of NUMA nodes involved:<\/p>\n<p><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-24272630\" src=\"https:\/\/www.europesays.com\/ie\/wp-content\/uploads\/2026\/02\/Arm_Scaling-llamacpp-On-Neoverse-N2-fig4.webp.png\" alt=\"\" width=\"954\" height=\"474\"  \/><\/p>\n<p>With the optimized ggml barrier, there is no obvious performance drop even if there are cross NUMA operations:<\/p>\n<p><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-24272631\" src=\"https:\/\/www.europesays.com\/ie\/wp-content\/uploads\/2026\/02\/Arm_Scaling-llamacpp-On-Neoverse-N2-fig5.webp.png\" alt=\"\" width=\"1284\" height=\"420\"  \/><\/p>\n<p>Mul Mat operation optimization<\/p>\n<p>The Mul Mat operation computation uses three tensor buffers: dst (for example, attn_out, ffn_gate, ffn_out, ffn_up, Kcur, Qcur, Vcur, FP32), src0 (weight), and src1 (for example, attn_norm, ffn_gate_par, ffn_norm, kqv_out FP32).<\/p>\n<p>These buffers are allocated using malloc which is not NUMA aware. As a result, many NUMA memory accesses occur when threads exceed a single NUMA node.<\/p>\n<p>The optimization splits the buffers into segments so that computation threads access memory from their local NUMA node whenever possible.<\/p>\n<p>For dst and src0 buffers, each thread accesses a portion of the buffers based on the thread ID. By splitting the buffer into N segments, where N is the number of NUMA nodes, each thread can access the portion of buffers in its local NUMA node if:<\/p>\n<ul>\n<li>Thread id is well mapped to the physical core id, which could be set through affinity.<\/li>\n<li>The segments are moved to the desired NUMA nodes, which could be done through move_pages() system call.<\/li>\n<\/ul>\n<p><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-24272632\" src=\"https:\/\/www.europesays.com\/ie\/wp-content\/uploads\/2026\/02\/Arm_Scaling-llamacpp-On-Neoverse-N2-fig6.webp.png\" alt=\"\" width=\"1280\" height=\"404\"  \/><\/p>\n<p>During the Mul Mat operation computation, src1 is quantized and stored in another buffer called wdata. The Mul Mat operation is then computed as:<\/p>\n<p>dst = src0 * wdata<\/p>\n<p>When src0 and dst is accessed by each threads according to thread id, wdata is a buffer which needs to be accessed entirely by those threads.<\/p>\n<p>The approach is based on the fact that quantization is not the main hotspot in the mul_mat operation. A NUMA-local wdata buffer is created for each NUMA node, and all threads in a NUMA node quantize src1 into its own wdata. As a result, there are N copies of wdata.<\/p>\n<p><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-24272633\" src=\"https:\/\/www.europesays.com\/ie\/wp-content\/uploads\/2026\/02\/Arm_Scaling-llamacpp-On-Neoverse-N2-fig7.webp.png\" alt=\"\" width=\"1282\" height=\"380\"  \/><\/p>\n<p>With the optimized tensor data layout for the Mul Mat operation computation, we can see there is a clear performance uplift:<\/p>\n<p><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-24272634\" src=\"https:\/\/www.europesays.com\/ie\/wp-content\/uploads\/2026\/02\/Arm_Scaling-llamacpp-On-Neoverse-N2-fig8.webp.png\" alt=\"\" width=\"1314\" height=\"512\"  \/><\/p>\n<p>Overall performance comparison<\/p>\n<p>We tested the llama.cpp batched benchmark and the results are:<\/p>\n<ul>\n<li>With the NUMA optimization there is a clear performance uplift.<\/li>\n<li>For S_TG t\/s, the best number of base is 26.52 with 32 threads. It is 41.15 with 40 threads for the NUMA optimization, which is a 55% uplift.<\/li>\n<li>For S t\/s, the best number of base is 48.73 with 36 threads. It is 74.67 with 54 threads for the NUMA optimization, which is a 53.2% uplift.<\/li>\n<\/ul>\n<p><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-24272635\" src=\"https:\/\/www.europesays.com\/ie\/wp-content\/uploads\/2026\/02\/Arm_Scaling-llamacpp-On-Neoverse-N2-fig9.webp.png\" alt=\"\" width=\"1258\" height=\"448\"  \/><\/p>\n<p>By capturing the memory bandwidth data of the NUMA optimization, we see the bandwidth balanced across NUMA nodes. Without the optimization, the bottleneck appears in NUMA node 0. This example uses a two-NUMA-node system:<\/p>\n<tr style=\"background-color: #ddd;\">\n<td style=\"padding: 5px;\"><strong>Use Cases<\/strong><\/td>\n<td style=\"padding: 5px;\"><strong>node 0 bandwidth GB\/s<\/strong><\/td>\n<td style=\"padding: 5px;\"><strong>node 1 bandwidth GB\/s<\/strong><\/td>\n<\/tr>\n<tr>\n<td style=\"padding: 5px;\">Stream test on numa node 0 with 64 threads<\/td>\n<td style=\"padding: 5px;\">119.8<\/td>\n<td style=\"padding: 5px;\">0<\/td>\n<\/tr>\n<tr>\n<td style=\"padding: 5px;\">llama.cpp numa node 0 with 64 threads<\/td>\n<td style=\"padding: 5px;\">104.9<\/td>\n<td style=\"padding: 5px;\">0<\/td>\n<\/tr>\n<tr>\n<td style=\"padding: 5px;\">llama.cpp numa node 0 and 1 with 128 threads<\/td>\n<td style=\"padding: 5px;\">70.6<\/td>\n<td style=\"padding: 5px;\">0.2<\/td>\n<\/tr>\n<tr>\n<td style=\"padding: 5px;\">optimized llama.cpp numa node 0 and 1 with 128 threads<\/td>\n<td style=\"padding: 5px;\">74.4<\/td>\n<td style=\"padding: 5px;\">72.2<\/td>\n<\/tr>\n<tr>\n<td style=\"padding: 5px;\">optimized llama.cpp numa node 0 and 1 with 54 threads<\/td>\n<td style=\"padding: 5px;\">97.4<\/td>\n<td style=\"padding: 5px;\">99.4<\/td>\n<\/tr>\n<p>Proof of concept\u00a0patch<\/p>\n<p>The proof-of-concept NUMA optimization patch was reviewed by a llama.cpp author, but was not merged as interest in server and cloud use cases is low. You can find the <a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/14232\" rel=\"nofollow noopener\" target=\"_blank\">NUMA optimization patch on GitHub<\/a>.<\/p>\n<p><\/p>\n<p>\t\t\t\t\t\t\tBolt Liu \u00a0\u00a0<a href=\"https:\/\/semiengineering.com\/author\/bolt-liu\/\" rel=\"nofollow noopener\" target=\"_blank\">(all posts)<\/a><\/p>\n<p><\/p>\n<p>\t\t\t\t\t\t\tBolt Liu is a software engineer at Arm China.<\/p>\n","protected":false},"excerpt":{"rendered":"This blog post explains the cross-NUMA memory access issue that occurs when you run llama.cpp in Neoverse. It&hellip;\n","protected":false},"author":2,"featured_media":333322,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","_share_on_mastodon":"0"},"categories":[74],"tags":[291,76015,18,19,17,23391,3589,159168,82],"class_list":["post-333321","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology","tag-ai","tag-arm","tag-eire","tag-ie","tag-ireland","tag-llama","tag-llms","tag-memory-access","tag-technology"],"share_on_mastodon":{"url":"https:\/\/pubeurope.com\/@ie\/116057010849870514","error":""},"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/posts\/333321","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/comments?post=333321"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/posts\/333321\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/media\/333322"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/media?parent=333321"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/categories?post=333321"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/tags?post=333321"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}