← Back to library
Adjacent technologyPeer reviewedACL 2024 paper

Catalog record / 12 Dec 2023

LLM in a flash established the flash-to-DRAM weight-loading baseline

The paper stores model parameters in flash and loads them into DRAM on demand, using windowing and row-column bundling to reduce transfers and make reads larger and more contiguous.

Why it matters

It is a foundational adjacent baseline for flash-backed LLM inference and a reminder that the phrase “LLM in a flash” describes a software-managed storage hierarchy, not High Bandwidth Flash.

Exact source title

LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Authors

Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, Karen Khatamifard, Minsik Cho, Carlo C Del Mundo, Mohammad Rastegari, Mehrdad Farajtabar

Tags

LLM in a flashweight loadingflash storagelimited DRAMACL 2024

Reading rule

This record preserves the source's maturity label. A roadmap, research result, or announced specification should not be treated as independent evidence of shipping hardware.