HBF researchPreprintResearch
Catalog record / 14 Aug 2026
DASH gives HBF two concurrent paths into the GPU
DASH combines a direct GPU–HBF route with an HBF–HBM–GPU relay and uses both concurrently to deliver MoE expert weights, while managing immutable experts separately from mutable KV-cache traffic.
Why it matters
It moves beyond treating HBF as passive capacity and directly studies expert placement, early reads, and dual UCIe data paths. The reported 1.94× throughput and 1.90× end-to-end speedup are simulator results from a preprint, not silicon measurements.
Exact source title
Beyond Capacity: Scalable MoE LLM Inference via High-Bandwidth Flash with Direct GPU and HBM Paths
Authors
Seeyeon Kim, Juhyeong Jin, Joo-Young Kim
Tags
Reading rule
This record preserves the source's maturity label. A roadmap, research result, or announced specification should not be treated as independent evidence of shipping hardware.