← Back to library
Adjacent technologyPreprintResearch

Catalog record / 8 Sep 2024

InstInfer makes attention and KV-cache management flash-aware

InstInfer offloads decoding attention and KV-cache data to computational storage drives, using flash-aware attention, KV management, and GPU-to-drive peer-to-peer transfers.

Why it matters

It is an early direct study of the long-context KV-cache path over flash, clarifying why internal flash bandwidth alone does not make an SSD-based design equivalent to package-local HBF.

Exact source title

InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference

Authors

Xiurui Pan, Endian Li, Qiao Li, Shengwen Liang, Yizhou Shan, Ke Zhou, Yingwei Luo, Xiaolin Wang, Jie Zhang

Tags

InstInferin-storage attentionKV cachecomputational storagelong context

Reading rule

This record preserves the source's maturity label. A roadmap, research result, or announced specification should not be treated as independent evidence of shipping hardware.