← Back to library
Adjacent technologyPreprintResearch

Catalog record / 3 Dec 2025

KVNAND puts both model weights and KV cache in 3D NAND

KVNAND proposes a DRAM-free architecture that stores model weights and KV cache in compute-enabled 3D NAND, with page-level KV mapping and head-group parallelism for long-context inference.

Why it matters

It directly covers the write-heavy KV-cache scenario that exposes HBF endurance and latency questions, but proposes in-flash computation rather than standards-track HBF.

Exact source title

KVNAND: Efficient On-Device Large Language Model Inference Using DRAM-Free In-Flash Computing

Authors

Lishuo Deng, Shaojie Xu, Jinwu Chen, Changwei Yan, Jiajie Wang, Zhe Jiang, Weiwei Shan

Tags

KVNANDKV cacheDRAM-freein-flash computinglong context

Reading rule

This record preserves the source's maturity label. A roadmap, research result, or announced specification should not be treated as independent evidence of shipping hardware.