Adjacent technologyPeer reviewedACL 2024 paper
Catalog record / 12 Dec 2023
LLM in a flash established the flash-to-DRAM weight-loading baseline
The paper stores model parameters in flash and loads them into DRAM on demand, using windowing and row-column bundling to reduce transfers and make reads larger and more contiguous.
Why it matters
It is a foundational adjacent baseline for flash-backed LLM inference and a reminder that the phrase “LLM in a flash” describes a software-managed storage hierarchy, not High Bandwidth Flash.
Exact source title
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Authors
Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, Karen Khatamifard, Minsik Cho, Carlo C Del Mundo, Mohammad Rastegari, Mehrdad Farajtabar
Tags
Reading rule
This record preserves the source's maturity label. A roadmap, research result, or announced specification should not be treated as independent evidence of shipping hardware.