Adjacent technologyPreprintResearch
Catalog record / 8 Sep 2024
InstInfer makes attention and KV-cache management flash-aware
InstInfer offloads decoding attention and KV-cache data to computational storage drives, using flash-aware attention, KV management, and GPU-to-drive peer-to-peer transfers.
Why it matters
It is an early direct study of the long-context KV-cache path over flash, clarifying why internal flash bandwidth alone does not make an SSD-based design equivalent to package-local HBF.
Exact source title
InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
Authors
Xiurui Pan, Endian Li, Qiao Li, Shengwen Liang, Yizhou Shan, Ke Zhou, Yingwei Luo, Xiaolin Wang, Jie Zhang
Tags
Reading rule
This record preserves the source's maturity label. A roadmap, research result, or announced specification should not be treated as independent evidence of shipping hardware.