← Back to library
Adjacent technologyPeer reviewedMICRO 2024 paper

Catalog record / 24 Sep 2024

Cambricon-LLM pairs an NPU with an in-flash-compute chiplet

Cambricon-LLM combines an NPU with a dedicated NAND chiplet using in-flash computing, on-die ECC, and hardware-aware tiling for large on-device models.

Why it matters

This peer-reviewed antecedent is useful for distinguishing HBF's high-bandwidth memory-tier model from architectures that add compute directly to a NAND chiplet.

Exact source title

Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLM

Authors

Zhongkai Yu, Shengwen Liang, Tianyun Ma, Yunke Cai, Ziyuan Nan, Di Huang, Xinkai Song, Yifan Hao, Jie Zhang, Tian Zhi, Yongwei Zhao, Zidong Du, Xing Hu, Qi Guo, Tianshi Chen

Tags

Cambricon-LLMNAND chipletin-flash computingon-deviceMICRO 2024

Reading rule

This record preserves the source's maturity label. A roadmap, research result, or announced specification should not be treated as independent evidence of shipping hardware.