Adjacent technologyPreprintResearch prototype
Catalog record / 22 Jan 2026
FlashMoE learns which inactive experts should stay off the SSD
FlashMoE keeps inactive experts on an SSD and uses a learned cache-replacement policy that combines recency and frequency, demonstrated on a user-grade desktop platform.
Why it matters
It covers the same sparse-expert residency problem as direct HBF MoE work, but through conventional SSD offload and software caching. That contrast helps isolate what package-local HBF bandwidth is intended to change.
Exact source title
FlashMoE: Reducing SSD I/O Bottlenecks via ML-Based Cache Replacement for Mixture-of-Experts Inference on Edge Devices
Authors
Byeongju Kim, Jungwan Lee, Donghyeon Han, Hoi-Jun Yoo, Sangyeob Kim
Tags
Reading rule
This record preserves the source's maturity label. A roadmap, research result, or announced specification should not be treated as independent evidence of shipping hardware.