← Back to library
Adjacent technologyPreprintResearch prototype

Catalog record / 22 Jan 2026

FlashMoE learns which inactive experts should stay off the SSD

FlashMoE keeps inactive experts on an SSD and uses a learned cache-replacement policy that combines recency and frequency, demonstrated on a user-grade desktop platform.

Why it matters

It covers the same sparse-expert residency problem as direct HBF MoE work, but through conventional SSD offload and software caching. That contrast helps isolate what package-local HBF bandwidth is intended to change.

Exact source title

FlashMoE: Reducing SSD I/O Bottlenecks via ML-Based Cache Replacement for Mixture-of-Experts Inference on Edge Devices

Authors

Byeongju Kim, Jungwan Lee, Donghyeon Han, Hoi-Jun Yoo, Sangyeob Kim

Tags

FlashMoEMoEexpert offloadSSD cacheedge inference

Reading rule

This record preserves the source's maturity label. A roadmap, research result, or announced specification should not be treated as independent evidence of shipping hardware.