{"id":123294,"date":"2026-08-31T18:04:55","date_gmt":"2026-08-31T18:04:55","guid":{"rendered":"https:\/\/www.europesays.com\/ch\/123294\/"},"modified":"2026-08-31T18:04:55","modified_gmt":"2026-08-31T18:04:55","slug":"workload-driven-hbf-substrate-for-capacity-scalable-llm-inference-huawei-eth-zurich-hust","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ch\/123294\/","title":{"rendered":"Workload-Driven HBF Substrate For Capacity-Scalable LLM Inference (Huawei, ETH Zurich, HUST)"},"content":{"rendered":"<p>Researchers at Huawei, ETH Z\u00fcrich, and HUST published a technical paper titled \u201cFLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference Acceleration.\u201d<\/p>\n<p>Abstract:<\/p>\n<p>\u201cLLM inference is increasingly constrained by accelerator memory capacity rather than compute throughput. This constraint is especially acute in single-accelerator and small-node inference systems, where limited on-package memory capacity restricts the size of deployable models. HBF is an emerging 3D-stacked NAND flash technology that provides multi-terabyte near-accelerator capacity, making it a promising capacity tier for storing LLM weights. However, existing HBF-based proposals face three adoption challenges: they (1) rely on coarse-grained static prefetching for LLM weights aiming to hide the microsecond-level read latency of the NAND flash device while maximizing HBF\u2019s read throughput, (2) expose NAND flash management tasks (e.g., refresh operations) to the accelerator-visible critical inference path, and (3) miss optimization opportunities to specialize and optimize the flash-management mechanisms to the workload behavior.<\/p>\n<p>Our goal is to design an efficient HBF substrate that integrates HBF as a memory-capacity tier alongside HBM while addressing these three challenges. To this end, we propose FLINT, a workload-driven HBF substrate for capacity-scalable LLM inference. FLINT introduces three mechanisms: (1) a hardware burst-buffer controller that dynamically coalesces and pipelines HBF reads aiming to utilize existing NAND flash buffers while sustaining high HBF bandwidth, (2) a phantom-plane refresh mechanism, which removes refresh from the critical inference path by moving refresh-related NAND flash operations outside the read foreground back via low-cost resource duplication, and (3) a read-only FTL, which replaces SSD-class support for arbitrary writes with a compact table that translates logical weight bursts to physical HBF locations.\u201d<\/p>\n<p><a href=\"https:\/\/arxiv.org\/abs\/2608.25062\" rel=\"nofollow noopener\" target=\"_blank\">Find the technical paper here<\/a>. August 2026.<\/p>\n<p>Oliveira, Geraldo F., Arash Tavakkol, Xiangyu Zhu, Ahmet Caner Y\u00fcz\u00fcg\u00fcler, Vamanan Arulchelvan, Lukas Cavigelli, Renzo Andri et al. \u201cFLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference Acceleration.\u201d arXiv preprint arXiv:2608.25062 (2026).<\/p>\n<p>\u00a0<\/p>\n<p>\u00a0<\/p>\n<p><\/p>\n","protected":false},"excerpt":{"rendered":"Researchers at Huawei, ETH Z\u00fcrich, and HUST published a technical paper titled \u201cFLINT: Efficiently Leveraging High Bandwidth Flash&hellip;\n","protected":false},"author":2,"featured_media":123295,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","_share_on_mastodon":"0"},"categories":[6],"tags":[60300,60301,50046,3740,60302,53907,60303,51875,60304,38415,60305,60306,28365,19685,60307,60308,60309,60310,60311,60312,48489,60313,51],"class_list":["post-123294","post","type-post","status-publish","format-standard","has-post-thumbnail","category-zurich","tag-3d-nand","tag-ai-accelerators","tag-data-movement","tag-eth-zurich","tag-flash-translation-layer","tag-flint","tag-ftl","tag-gpu","tag-hbf","tag-hbm","tag-heterogeneous-memory","tag-high-bandwidth-flash","tag-high-bandwidth-memory","tag-huawei","tag-huazhong-university-of-science-and-technology","tag-hust","tag-llm-inference","tag-memory-architecture","tag-memory-bandwidth","tag-memory-capacity","tag-nand-flash","tag-prefetching","tag-zurich"],"share_on_mastodon":{"url":"https:\/\/pubeurope.com\/@ch\/117191450786259691","error":""},"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ch\/wp-json\/wp\/v2\/posts\/123294","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ch\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ch\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ch\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ch\/wp-json\/wp\/v2\/comments?post=123294"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ch\/wp-json\/wp\/v2\/posts\/123294\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ch\/wp-json\/wp\/v2\/media\/123295"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ch\/wp-json\/wp\/v2\/media?parent=123294"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ch\/wp-json\/wp\/v2\/categories?post=123294"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ch\/wp-json\/wp\/v2\/tags?post=123294"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}