Adjacent technologyPreprintResearch
Catalog record / 3 Dec 2025
KVNAND puts both model weights and KV cache in 3D NAND
KVNAND proposes a DRAM-free architecture that stores model weights and KV cache in compute-enabled 3D NAND, with page-level KV mapping and head-group parallelism for long-context inference.
Why it matters
It directly covers the write-heavy KV-cache scenario that exposes HBF endurance and latency questions, but proposes in-flash computation rather than standards-track HBF.
Exact source title
KVNAND: Efficient On-Device Large Language Model Inference Using DRAM-Free In-Flash Computing
Authors
Lishuo Deng, Shaojie Xu, Jinwu Chen, Changwei Yan, Jiajie Wang, Zhe Jiang, Weiwei Shan
Tags
Reading rule
This record preserves the source's maturity label. A roadmap, research result, or announced specification should not be treated as independent evidence of shipping hardware.